~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #claude-code $ grep

#Claude Code

17 posts
13 #openai · #multi-agent
GPT-5.6 in Codex: Ultra mode buys 3 points, and Terra burns 2.6x more tokens to deliver worse results

GPT-5.6 has been in Codex since July 9, across all three tiers. Ultra mode coordinates four subagents that cooperate during the task and buys 3 points on Terminal-Bench (88.8% to 91.9%), at a much higher token cost. And CodeRabbit's independent test kills the reflex of picking the cheap tier: Terra burned 2.6x more tokens than Sol to pass 40.7% of tasks versus 63.7%. Price per token is not price per task.

15 Jul · 11 min ›
14 #openai · #anthropic
Codex Usage Limit: OpenAI Drops the 5-Hour Cap and Anthropic Answers Within Hours

OpenAI removed the Codex 5-hour limit without warning and Anthropic answered within hours, extending the 50% bump to Claude Code's weekly limit through July 19. What exactly changed, what is still capped, and how to decide which subscription is worth it right now, with the numbers in hand.

13 Jul · 8 min ›
15 #produto-ia · #llm
Claude Sonnet 5: How to Read the Benchmarks in Practice

Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.

30 Jun · 9 min ›
16 #openai · #ia
Claude Code vs Codex: Codex Wins the Terminal by 13 Points, Claude Wins the Hard Repo by 10

Claude Code or Codex? The answer comes with numbers: Codex opens a 13-point lead on Terminal-Bench (82.7% vs 69.4%) and Claude Code opens a 10-point lead on SWE-bench Pro (69.2% vs 58.6%), which is the benchmark for real multi-file problems. On SWE-bench Verified they tie. Here is the verdict by scenario, the real cost per dev, and the criterion that matters more than quality: how much control you want during the task.

19 Jun · 10 min ›
17 #tool-use · #llm-as-a-judge
Code Review Prompt: 5 Patterns That Raise Signal from 12% to 67%

An LLM code review prompt your team will actually read: five patterns (diff-anchored, severity gate, tool use before guessing, mandatory citation, and self-grading with a threshold) that raise the signal ratio from ~20% to above 60%. Comes with a complete REVIEW.md and a GitHub Actions workflow ready to plug into a Laravel project.

24 May · 15 min ›
Meet the Clã Beer and Code
playing