~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / author / lucas-souza-virgu $ whoami
Lucas Souza

Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

78 posts
61 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
62 #llm · #modelos-open-source
Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

Moonshot AI published the Kimi K3 open weights a day ahead of schedule: 2.8 trillion parameters, a modified MIT license and native MXFP4 format. The math the hype skips: 1.4 TB of VRAM just to load the weights, an 8-GPU node that doesn't add up, and a break-even of 6,200 tokens per second to match the API at $3 / $15. Plus what US sanctions can actually reach when you download weights from Brazil.

27 Jul · 14 min ›
63 #openai · #boas-praticas
10 Codex Best Practices (Straight from the Official Documentation)

We combed through the new official Codex documentation (learn.chatgpt.com) and boiled it down to the 10 best practices OpenAI itself recommends: outcome-first prompting, a living AGENTS.md, restrictive permissions, model selection, skills, cloud, and automation.

26 Jul · 8 min ›
64 #claude · #anthropic
Fable 5 free credits: claiming them turns on usage billing on your account without warning

Anthropic handed out $100 (about R$ 540) in free Fable 5 credits to Pro and Team plans, but claiming them requires a card and leaves extra usage enabled: when the credit runs out, your account starts getting billed for usage, with no separate warning. A step-by-step to check whether billing is turned on in your account and how to turn it off, or how to use the $100 with a spending cap and auto-reload disabled.

23 Jul · 8 min ›
65 #openai · #noticias
GPT-6 Release Timeline: From Altman in Washington to the Astra Launch (Tracker Closed)

In July, Bloomberg reported Altman's briefing to the US government and the internet read it as an imminent GPT-6 launch. This tracker logged, with a date and a source, what was confirmed and what was speculation until the model shipped on September 3, 2026 as GPT-6 Astra. It stays up as a timeline; the numbers, the pricing and the comparison with Fable 5.1 and GPT-5.6 Sol are in the comparison post.

22 Jul · 7 min ›
66 #ia · #llm
How to Run an LLM Locally: Step by Step with Ollama in 2 Commands (No GPU)

A step-by-step guide to running an LLM locally with Ollama: install it, run qwen3:8b in 2 commands, and plug it into your code through the OpenAI-compatible endpoint. It runs on 8 GB of RAM with no GPU required. As a bonus, the VRAM math by model size and when local beats the API.

22 Jul · 9 min ›
67 #ia · #claude
Fable 5 and the Jacobian Conjecture: the 90-year-old problem AI appears to have knocked down

Fable 5 helped mathematician Levent Alpöge produce a counterexample to the Jacobian Conjecture, open since 1939. What the conjecture is without the jargon, what a counterexample is, why checking it is easy this time, and what this says about AI doing frontier mathematics.

20 Jul · 9 min ›
68 #ia · #llm
Qwen 3.8 Max (2.4T): Benchmark vs Kimi K3, Pricing and Open Weights

Alibaba dropped Qwen 3.8 Max, a 2.4-trillion-parameter MoE it claims is the second-best model in the world, behind only Fable 5. With no public benchmark at launch, the only independent test scored it 80/100 against Kimi K3's 83. Here: what it is, what it costs (Token Plan from $6 to $68), how to access it via API, and where each model in the Qwen 3.8 family fits.

20 Jul · 9 min ›
69 #ia · #llm
Kimi K3 Pricing: $3/$15 per Million Tokens, Is It Free? Benchmarks vs Claude

Kimi K3 costs $3 input and $15 output per million tokens on the API (the chat at kimi.com is free), 1.7x cheaper than Claude Opus 4.8, not the "5x" your timeline is claiming. We did the honest pricing math, separated the verifiable benchmarks from the hype, and show where it beats Claude (and where it doesn't).

17 Jul · 10 min ›
70 #openai · #multi-agent
GPT-5.6 in Codex: Ultra mode buys 3 points, and Terra burns 2.6x more tokens to deliver worse results

GPT-5.6 has been in Codex since July 9, across all three tiers. Ultra mode coordinates four subagents that cooperate during the task and buys 3 points on Terminal-Bench (88.8% to 91.9%), at a much higher token cost. And CodeRabbit's independent test kills the reflex of picking the cheap tier: Terra burned 2.6x more tokens than Sol to pass 40.7% of tasks versus 63.7%. Price per token is not price per task.

15 Jul · 11 min ›
71 #openai · #anthropic
Codex Usage Limit: OpenAI Drops the 5-Hour Cap and Anthropic Answers Within Hours

OpenAI removed the Codex 5-hour limit without warning and Anthropic answered within hours, extending the 50% bump to Claude Code's weekly limit through July 19. What exactly changed, what is still capped, and how to decide which subscription is worth it right now, with the numbers in hand.

13 Jul · 8 min ›
72 #produto-ia · #llm
Claude Sonnet 5: How to Read the Benchmarks in Practice

Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.

30 Jun · 9 min ›
Meet the Clã Beer and Code
playing