#Claude
Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.
Anthropic handed out $100 (about R$ 540) in free Fable 5 credits to Pro and Team plans, but claiming them requires a card and leaves extra usage enabled: when the credit runs out, your account starts getting billed for usage, with no separate warning. A step-by-step to check whether billing is turned on in your account and how to turn it off, or how to use the $100 with a spending cap and auto-reload disabled.
Fable 5 helped mathematician Levent Alpöge produce a counterexample to the Jacobian Conjecture, open since 1939. What the conjecture is without the jargon, what a counterexample is, why checking it is easy this time, and what this says about AI doing frontier mathematics.
Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.
Claude Code or Codex? The answer comes with numbers: Codex opens a 13-point lead on Terminal-Bench (82.7% vs 69.4%) and Claude Code opens a 10-point lead on SWE-bench Pro (69.2% vs 58.6%), which is the benchmark for real multi-file problems. On SWE-bench Verified they tie. Here is the verdict by scenario, the real cost per dev, and the criterion that matters more than quality: how much control you want during the task.
Fable 5 or Opus 4.8, which one should you use? A straight verdict by task type, with the 10 concrete situations where Fable gets it done and Opus 4.8 did it badly or not at all: migration at scale, code from a screenshot, long-running agents and reasoning over documents. Plus what each one costs and the cases where Opus 4.8 is still the right call.
An LLM code review prompt your team will actually read: five patterns (diff-anchored, severity gate, tool use before guessing, mandatory citation, and self-grading with a threshold) that raise the signal ratio from ~20% to above 60%. Comes with a complete REVIEW.md and a GitHub Actions workflow ready to plug into a Laravel project.