~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / claude-opus-5-got-worse $
News

Did Claude Opus 5 Get Worse? Better on Benchmarks, Worse to Live With

LS Lucas Souza · · 7 min read
Did Claude Opus 5 Get Worse? Better on Benchmarks, Worse to Live With

There's a thread on r/ClaudeAI with 537 upvotes titled "I can't use Opus 5 anymore". There's another one, with 164 upvotes, asking whether Claude got worse over the last ten days.

Look at the comment counts: 187 on the first, 158 on the second.

That's almost one comment per vote. And that number says more than the title does.

A thread with lots of votes and few comments is consensus — everybody agrees and moves on. A thread where comments nearly match votes is disagreement. Half the crowd saying the model went off the rails, the other half saying it's great.

When both sides are that big, it's usually not one of them lying. It's that both things are true at the same time, measured with different rulers.

And that's what's happening with Claude Opus 5. It didn't get worse on benchmarks: it got worse to live with. Both readings are right, and the explanation has a date and a source.

Did Claude Opus 5 get worse? The verdict that shows up in every review

Opus 5 shipped on July 24, 2026, at the same price as Opus 4.8: $5 per million input tokens, $25 on output.

On benchmarks, it went up. As something to live with, it soured.

MindStudio's analysis sums up the pattern that shows up in practically every review since launch: it's the best model a lot of people have ever used, and the one those same people like using the least.

Google has already baked this in. If you search "did claude opus 5 get worse" today, the AI Overview opens with a "yes" and lists three complaints:

  • Excessive verbosity. Answers that run way too long, with features nobody asked for.
  • Worse code. Poorer judgment than Opus 4.8 on simple tasks.
  • Too much caution. Asks for confirmation too often, avoids making the call.

When the search engine itself is assembling the answer, the phenomenon has stopped being an anecdote. But "the phenomenon exists" and "the model regressed" are different claims. And the second one hasn't been proven.

What you can measure and what's perception

Pull the three complaints apart, because they don't carry the same technical weight.

Verbosity is measurable. Count output tokens on the same task, same prompt, different model. If Opus 5 spends more to deliver the same thing, it shows up on the bill. It's objective, and you can check it today in your own history.

Over-engineering is semi-measurable. You can count files touched, lines added and functions created outside the requested scope. It's not a clean metric, but it's a lot better than "felt wordy to me".

"Worse judgment" is perception. And it's the slipperiest of the three, because it runs into two well-known biases: you switched models and also switched the way you prompt, and you remember the old model's wins better than its misses.

It's not that the complaint is false. It's that it's the only one nobody has managed to instrument so far — and it's exactly the one that shows up most in the threads.

There's a documented mechanism. And it's not the model.

Here's the part almost nobody is connecting.

In July, Anthropic cut more than 80% of Claude Code's built-in system prompt for the Claude 5 generation of models. With no measurable loss on coding evals, according to the company itself.

Now think about what used to live in that system prompt.

An instruction to be concise. An instruction not to create files nobody asked for. An instruction not to explain the obvious. An instruction to ask before doing anything destructive. All that behavioral scaffolding that made the model look polite by default.

Take the scaffolding away and what's left is the raw model.

A more capable raw model is going to do more. It's going to write more. It's going to anticipate more. It's going to propose the abstraction you didn't ask for — because technically it even makes sense, and there's nobody left saying "don't do that without being asked".

In other words: part of what the community is reading as "the model got worse" may be "the defaults disappeared". Responsibility for behavior moved from Anthropic's system prompt to your CLAUDE.md. If you had a lean, explicit file, you barely felt it. If you were leaning on the defaults, you watched the behavior change overnight without understanding why.

That's a hypothesis, not a confirmed fact — Anthropic hasn't publicly tied one to the other. But it's the hypothesis with more documentary evidence than any "they nerfed the model to save GPU" theory, which shows up in every thread and never comes with a source.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

And then there's the fallback, which nobody remembers to check

Another documented detail that muddies perception: Anthropic maintains a support article called Why Claude switched models in your conversation with Opus 5.

There is a fallback. Under certain conditions, your conversation switches models.

So when someone says "Opus 5 is dumb today", there's a real chance that in that session they weren't talking to Opus 5 the whole time. And since it happens silently, the perception of inconsistency grows without any change to the model at all.

That doesn't invalidate the complaints. But it explains why the experience varies so much between two people using "the same" model on the same day.

Which is better, Claude Opus or Sonnet?

A question that shows up right in the People Also Ask for this search, and one that gets more interesting with this context.

If the main complaint against Opus 5 is that it does too much, then for every task where "doing too much" is the risk — CRUD, a targeted tweak, a mechanical refactor, writing a test for a known case — Sonnet 5 is the better choice, not the cheap choice. Less spare capacity is less spare capacity to make things up with.

Save Opus for the problem that actually demands it: a multi-file bug, an architecture decision, legacy code with no tests. That's where the extra capacity pays off.

We've already compared Claude Code and Codex with numbers on the table, and the conclusion holds here too: the useful question is almost never "which model is best". It's "how much control do I want during the task".

How to tame it without switching models

If the system prompt hypothesis is right, the fix isn't waiting on Anthropic. It's putting the guardrails that disappeared back into your CLAUDE.md.

Here's what usually takes care of most of the verbosity:

## Escopo
- Faça exatamente o que foi pedido. Nada além.
- Não crie arquivo novo a menos que seja necessário para a tarefa.
- Não escreva documentação, README ou comentário a menos que eu peça.

## Resposta
- Responda em no máximo 4 linhas, exceto quando eu pedir detalhe.
- Sem preâmbulo e sem resumo do que você acabou de fazer.
- Nada de "ótima pergunta" ou variações.

## Antes de agir
- Mudança destrutiva ou fora do escopo: pergunte antes.

It's not magic and it's not new — it's the same principle behind the CLAUDE.md clauses we've already broken down. What changed is the urgency: these rules used to be an optimization. Now, with the defaults cut, they're what holds the behavior in place.

And before you pick a fight with the model, run the two-minute test: take a task you already ran on Opus 4.8, run it the same way on Opus 5, and compare output tokens and files touched. If the difference is in the output and not in the correctness, your problem is scope, not intelligence. That gets fixed in the context file.

The summary

Opus 5 probably didn't regress. It got more capable and less restrained, and nobody put back the restraint Anthropic took away.

That leaves three things standing:

  • The community's complaint is real, and it's not a tantrum. The behavior really did change.
  • The most likely explanation is documented and boring: the defaults left the system prompt. No conspiracy theory about saving GPU required.
  • The fix is on your side, and it's cheaper than switching vendors.

And here's the lesson that outlives this model and the next one: when the agent's behavior changes without you touching anything, the first place to look isn't the model. It's what the model is reading before it answers you.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing