#Contexto
The threads say Claude Opus 5 regressed, and Google already answers yes. But the most likely explanation isn't a model nerf: Anthropic cut more than 80% of Claude Code's built-in system prompt for the Claude 5 generation. The restraint defaults are gone, and the responsibility moved to your CLAUDE.md. What you can measure, what's perception, and how to tame it without switching vendors.
The operational fundamentals of LLMs with no parrot analogy: what the model does at each token, why it's stateless, how the context window degrades long before the limit, and why temperature 0 isn't deterministic. Each concept closes with the architectural consequence it forces you to build, with the official Anthropic and OpenAI docs, the Lost in the Middle paper, and Chroma's context rot study in hand.
Half the prompt techniques disappeared because the model learned them on its own. The other half became API parameters. What survived, what turned into folklore, and why in a real system the problem stopped being the sentence and became what goes into the context window.
You send a one-line question and /usage reports a whole day's worth of consumption. Saving tokens in a coding assistant has nothing to do with prompt size: it's about prefix caching. How it works in Claude Code, Codex, and Cursor, the seven actions that invalidate it without you noticing, how to measure it with cache_read vs cache_creation, and eight levers to stretch the session.