~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #claude $ grep

#Claude

19 posts
13 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
14 #claude · #anthropic
Fable 5 free credits: claiming them turns on usage billing on your account without warning

Anthropic handed out $100 (about R$ 540) in free Fable 5 credits to Pro and Team plans, but claiming them requires a card and leaves extra usage enabled: when the credit runs out, your account starts getting billed for usage, with no separate warning. A step-by-step to check whether billing is turned on in your account and how to turn it off, or how to use the $100 with a spending cap and auto-reload disabled.

23 Jul · 8 min ›
15 #ia · #claude
Fable 5 and the Jacobian Conjecture: the 90-year-old problem AI appears to have knocked down

Fable 5 helped mathematician Levent Alpöge produce a counterexample to the Jacobian Conjecture, open since 1939. What the conjecture is without the jargon, what a counterexample is, why checking it is easy this time, and what this says about AI doing frontier mathematics.

20 Jul · 9 min ›
16 #produto-ia · #llm
Claude Sonnet 5: How to Read the Benchmarks in Practice

Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.

30 Jun · 9 min ›
17 #openai · #ia
Claude Code vs Codex: Codex Wins the Terminal by 13 Points, Claude Wins the Hard Repo by 10

Claude Code or Codex? The answer comes with numbers: Codex opens a 13-point lead on Terminal-Bench (82.7% vs 69.4%) and Claude Code opens a 10-point lead on SWE-bench Pro (69.2% vs 58.6%), which is the benchmark for real multi-file problems. On SWE-bench Verified they tie. Here is the verdict by scenario, the real cost per dev, and the criterion that matters more than quality: how much control you want during the task.

19 Jun · 10 min ›
18 #ia · #ai-agents
Fable 5 vs Opus 4.8: which one to use (and the 10 tasks where the difference shows)

Fable 5 or Opus 4.8, which one should you use? A straight verdict by task type, with the 10 concrete situations where Fable gets it done and Opus 4.8 did it badly or not at all: migration at scale, code from a screenshot, long-running agents and reasoning over documents. Plus what each one costs and the cases where Opus 4.8 is still the right call.

09 Jun · 11 min ›
19 #tool-use · #llm-as-a-judge
Code Review Prompt: 5 Patterns That Raise Signal from 12% to 67%

An LLM code review prompt your team will actually read: five patterns (diff-anchored, severity gate, tool use before guessing, mandatory citation, and self-grading with a threshold) that raise the signal ratio from ~20% to above 60%. Comes with a complete REVIEW.md and a GitHub Actions workflow ready to plug into a Laravel project.

24 May · 15 min ›
Meet the Clã Beer and Code
playing