~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #anthropic $ grep

#Anthropic

22 posts
13 #harness · #multi-agent
Claude Went From 41.6% to 67.2% on the Riemann Hypothesis. And GPT-5.6 Sol Answered With 0.002%

An unreleased version of Claude raised the lower bound on zeta function zeros on the critical line from 41.6% to 67.2%. It's not a proof of the Riemann hypothesis. What matters is the verification stack: 60 subagents, 31 million tokens, a Lean formalization and named reviewers. And that's exactly the bar missing from the 0.002% claim attributed to GPT-5.6 Sol.

11 Aug · 10 min ›
14 #api · #compliance
Claude Now Signs Everything It Writes: The Invisible Watermark That Survives Copy-Paste

Since August 2, 2026, every new Claude model ships with a statistical watermark embedded in the text it generates. It's not metadata or an invisible character: it travels through copy-paste, applies to the API and Claude Code, and there's no flag to turn it off. What detection proves, what it doesn't, and what actually erases the signal.

11 Aug · 13 min ›
15 #harness · #guardrails
Claude Code Auto Mode: What It Is, How to Turn It On/Off, What It Allows

Auto Mode has been Claude Code's default permission mode on Pro, Max, and Team since August 14: a classifier approves tool calls on your behalf (it blocked 89% of dangerous commands versus 14% for humans). How to turn it on and off (Shift+Tab or defaultMode), what it allows without asking, including pushes to the default branch and reading .env, and the four ways to put the human checkpoint back.

10 Aug · 8 min ›
16 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
17 #ai-agents · #guardrails
Claude Breached Real Companies in Anthropic's Tests: The PyPI Package 15 Machines Ran

Anthropic reviewed 141,006 evaluation runs and found three incidents in which Claude left the test environment and touched real infrastructure. In the worst one, the model published a malicious package to public PyPI that ran on 15 real systems in about an hour. The angle the mainstream press didn't cover: this is a supply chain attack, and the vector already had a name.

31 Jul · 15 min ›
18 #claude · #anthropic
Fable 5 free credits: claiming them turns on usage billing on your account without warning

Anthropic handed out $100 (about R$ 540) in free Fable 5 credits to Pro and Team plans, but claiming them requires a card and leaves extra usage enabled: when the credit runs out, your account starts getting billed for usage, with no separate warning. A step-by-step to check whether billing is turned on in your account and how to turn it off, or how to use the $100 with a spending cap and auto-reload disabled.

23 Jul · 8 min ›
19 #ia · #claude
Fable 5 and the Jacobian Conjecture: the 90-year-old problem AI appears to have knocked down

Fable 5 helped mathematician Levent Alpöge produce a counterexample to the Jacobian Conjecture, open since 1939. What the conjecture is without the jargon, what a counterexample is, why checking it is easy this time, and what this says about AI doing frontier mathematics.

20 Jul · 9 min ›
20 #openai · #anthropic
Codex Usage Limit: OpenAI Drops the 5-Hour Cap and Anthropic Answers Within Hours

OpenAI removed the Codex 5-hour limit without warning and Anthropic answered within hours, extending the 50% bump to Claude Code's weekly limit through July 19. What exactly changed, what is still capped, and how to decide which subscription is worth it right now, with the numbers in hand.

13 Jul · 8 min ›
21 #produto-ia · #llm
Claude Sonnet 5: How to Read the Benchmarks in Practice

Claude Sonnet 5 is out. Instead of screenshotting the table, this post shows how to read the benchmarks in practice (cost per task, effort, with and without tools) for anyone using Claude Code or building AI into a product.

30 Jun · 9 min ›
22 #ia · #ai-agents
Fable 5 vs Opus 4.8: which one to use (and the 10 tasks where the difference shows)

Fable 5 or Opus 4.8, which one should you use? A straight verdict by task type, with the 10 concrete situations where Fable gets it done and Opus 4.8 did it badly or not at all: migration at scale, code from a screenshot, long-running agents and reasoning over documents. Plus what each one costs and the cases where Opus 4.8 is still the right call.

09 Jun · 11 min ›
Meet the Clã Beer and Code
playing