How to Save Tokens in Claude Code, Codex, and Cursor: Stop Breaking the Cache
You open Claude Code in the morning and work all day in the same session. At six in the evening you ask a one-line question — "is this test passing?" — and /usage reports consumption as if you had sent the entire repository.
You did. Every single time.
Saving tokens in a coding assistant has very little to do with the size of what you type. The API keeps nothing between requests: every turn resends the system prompt, the CLAUDE.md, every previous message, every tool result, and only then your question. What decides your bill at the end of the month is how much of that the server can reuse from the previous turn.
In this post: how prompt caching works in the three harnesses devs actually use (Claude Code, Codex, and Cursor), the seven actions that break that cache without you noticing, how to measure whether it's working, and eight concrete levers to stretch your session without blowing through the limit.
TL;DR
- What it is: how to save tokens in coding assistants, why the session gets expensive, and what to do about it.
- Tools: Claude Code, OpenAI Codex, Cursor.
- Cost/Access: nothing to install. It's commands, configuration, and habit.
- The rule that sums it all up: caching is prefix comparison. A change anywhere in the beginning recomputes everything that comes after it.
- Useful link: How Claude Code uses prompt caching, the most explicit documentation on invalidation that exists today.
Tokens aren't what you type
Prompt caching works by exact prefix matching. The server compares the beginning of your request with what it processed recently and reuses the identical part. There is no per-file or per-segment cache: the match is exact, and a change anywhere in the prefix recomputes everything that comes after it.
That's why Claude Code orders every request in layers, from what changes least to what changes most:
| Layer | Content | Changes when |
|---|---|---|
| System prompt | Core instructions, tool definitions, output style | The set of loaded tools changes, or Claude Code is updated |
| Project context | CLAUDE.md, automatic memory, rules | The session starts, or after /clear and /compact |
| Conversation | Your messages, responses, tool results | Every turn |
A change in the conversation is cheap: the two layers above it stay cached. A change in the system prompt is expensive: everything after it now sits behind a different prefix and gets reprocessed.
And the price difference is big. On the Anthropic API, a cache read costs 0.1× the normal input price — 10%. A write costs 1.25× (or 2× if you opt into the one-hour TTL). In other words: a paid cache write pays for itself after a single read on the five-minute TTL, and after two on the one-hour TTL.
This is not an implementation detail. The team that builds Claude Code treats cache hit rate like infrastructure uptime and declares an emergency when it drops — plan mode, deferred tool loading, and compaction were all designed around this constraint. If the tool's harness is designed that way, the way you use it should be too. It's this kind of discipline — measure, understand the mechanism, change the habit — that we practice live every week in the Clã Beer and Code: it's a subscription, it's paid, and it's exactly the environment this post describes.
At OpenAI the mechanics are the same, with their own numbers: the cache requires a prefix of at least 1,024 tokens, the official recommendation is static content at the beginning, dynamic content at the end, and as of GPT-5.6 the prompt_cache_key became necessary for reliable matching (each key handles roughly 15 requests per minute).
The seven actions that burn your cache
These cause one slow, expensive turn. Just once — after that the new prefix is cached. The problem is that most of them look free in the moment.
| Action | Why |
|---|---|
/model |
Each model has its own cache. Switching means rereading the entire conversation with zero hits |
/effort |
The effort level is also part of the cache key. Claude Code asks for confirmation before applying it |
| Turning on fast mode | Adds a header that is part of the key. It costs once per conversation — cheap at the start of the session, expensive at the end |
| Connecting/disconnecting an MCP | Tool definitions live in the system prompt. It only skips invalidation when the tools are deferred by tool search (the default on supported models) |
| Deny rule for an entire tool | Bash or WebFetch as a deny rule removes the tool from context. A scoped rule (Bash(rm *)) doesn't touch the prefix |
/compact |
Replaces the history with a summary. By definition it shares no prefix with what came before |
| Updating Claude Code | A new version usually changes the system prompt or tool definitions |
The gotcha that catches the most people: opusplan. With this setting the model resolves to Opus in plan mode and Sonnet in execution — so every plan mode toggle is a model switch, with a fresh cache. You think you're just changing modes.
Now the other side, which almost nobody knows. These do not break the cache:
- Editing files in the repository. File content only enters context when Claude reads it. Editing afterward doesn't rewrite the old read: the harness appends a notice that the file changed.
- Invoking skills and commands. They inject instructions as a user message at the point of invocation. Nothing before it changes.
/rewind. Truncates the conversation to an earlier point — which is exactly the content the cache was built from. You land back on a prefix that's already warm.- Switching permission mode. It changes neither the system prompt nor the tools. The exception is plan mode with
opusplan, for the reason above. - Subagent. It starts its own conversation, with its own cache. On the parent's side, the call and the result are just appended to the end.
And two that don't break the cache and don't apply the change either: editing the CLAUDE.md mid-session and switching output style. Both are read once at the start and kept in memory. The edit only kicks in at the next /clear, /compact, or restart. If you've ever edited the CLAUDE.md assuming Claude would start obeying on the next message — it didn't.
Measure before you optimize
Every API response carries two counters:
| Field | Meaning |
|---|---|
cache_creation_input_tokens |
Tokens written to the cache this turn, billed at 1.25× |
cache_read_input_tokens |
Tokens served from the cache, billed at ~10% of normal input |
A high read-to-write ratio means it's working. If creation stays high turn after turn, something in your prefix is changing.
You can track this live in a statusline:
#!/usr/bin/env bash
# ~/.claude/statusline.sh — registre com /statusline. Requer jq.
input=$(cat)
read_tokens=$(echo "$input" | jq -r '.current_usage.cache_read_input_tokens // 0')
write_tokens=$(echo "$input" | jq -r '.current_usage.cache_creation_input_tokens // 0')
if [ "$write_tokens" -gt 0 ]; then
ratio=$(awk "BEGIN { printf \"%.1f\", $read_tokens / $write_tokens }")
else
ratio="--"
fi
printf 'cache %s lidos / %s escritos (%sx)' "$read_tokens" "$write_tokens" "$ratio"
On top of that:
/usageshows, on a paid plan, usage attribution by skill, subagent, plugin, and MCP server — each as a percentage of the total. It also flags behaviors that account for 10% or more of recent usage, like long context or cache misses./contextshows what is taking up space right now.- In Cursor, the Context Usage Report breaks usage down by system prompt, tool definitions, rules, skills, MCPs, and subagents. Just click the agent's context ring.
- On the OpenAI API, the field is
cached_tokensinsideusage.input_tokens_details(Responses) orusage.prompt_tokens_details(Chat Completions).
If cached_tokens stays at zero on requests with an identical prefix, you have a silent invalidator: a datetime.now() in the system prompt, a per-request UUID, a json.dumps without sort_keys. It's the same family of problem as the most common token leaks in production, just inside your editor.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãEight levers to save tokens and stretch the session
1. /clear costs zero. /compact is one big request.
This flips a lot of people's intuition. To produce the summary, /compact sends a separate request with the same system prompt, the same tools, and the same history, plus the instruction to summarize. With a warm cache it reads the prefix from the cache and costs a fraction. After a long break, with no cache, it reprocesses the entire history as uncached input.
In other words: /compact is most expensive precisely when you pick an old session back up. /clear, on the other hand, sends no request at all.
Rule of thumb: want continuity, /compact at a natural break between tasks. Want a fresh start, /clear. Want to abandon a path, /rewind — which goes back to an already-cached prefix instead of building a new one.
2. Trim the CLAUDE.md
It gets loaded into context at the start of every session. If it has detailed instructions for PR review and database migrations, those tokens are there even when you're tweaking CSS. The official recommendation is to keep it under 200 lines, essentials only, and move specialized workflows into skills — which load on demand.
3. Filter the output before it becomes context
A PreToolUse hook can rewrite the command before Claude sees the result. Instead of it reading a 10,000-line log to find the error, the hook does the grep and returns the lines that matter — from tens of thousands of tokens down to hundreds.
#!/usr/bin/env bash
# ~/.claude/hooks/filtra-teste.sh — só as falhas voltam para o contexto.
input=$(cat)
cmd=$(echo "$input" | jq -r '.tool_input.command')
if [[ "$cmd" =~ ^(npm\ test|pytest|php\ artisan\ test|go\ test) ]]; then
novo="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"
echo "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"allow\",\"updatedInput\":{\"command\":\"$novo\"}}}"
else
echo "{}"
fi
If hooks are still new territory for you, there's a whole post on the subject in Hooks, Slash Commands, and MCPs: the anatomy of a productive harness.
4. Prefer the CLI over the MCP when the CLI exists
gh, aws, gcloud, sentry-cli are more context-efficient than the equivalent MCPs, because they don't add a per-tool listing. And run /mcp to turn off any server you aren't using.
The size of the problem has a number: Cursor moved to dynamic discovery of MCP tools — only the names go into the static context, and the agent fetches the schema when it needs it — and measured, in an A/B test, a 46.9% reduction in total agent tokens on runs that called an MCP tool.
Almost half. Just from not loading schemas that won't be used.
5. Delegate verbose operations to a subagent
Running a test suite, fetching documentation, processing a log. The output stays in the subagent's context and only the summary comes back to the main conversation. Important detail: a subagent uses the five-minute TTL even on a subscription — the one-hour TTL applies only to the main conversation.
6. Pin model and effort per phase of work
This isn't "always use the cheap model." It's about not switching back and forth. Every switch in the middle of a long task pays for the full reprocessing, and the gain from the cheaper model can be smaller than the cache loss. Pick the model and effort at the start of the session; if you need to change, change at a point where you were going to clear the context anyway.
Sonnet handles most coding work. Opus is for architecture decisions and multi-step reasoning. For a simple subagent, you can specify model: haiku in the configuration.
7. On the API, a stable, versioned prefix
If you're the one building the harness, the same rule applies: frozen system prompt, deterministic tool order (sort by name), no timestamp or request ID at the beginning of the prompt. Serialize JSON with sorted keys. Dynamic context goes at the end, after the last breakpoint. This is the same lever we already covered from the builder's side in Cutting costs by 80%: prompt caching, batch, and when NOT to use a reranker.
In Codex, the equivalent controls are configuration: model_context_window and model_auto_compact_token_limit define when compaction fires, and the compaction API allows server-side compaction via context_management with compact_threshold — no separate call.
8. Understand your TTL
In Claude Code on a subscription, the one-hour TTL is requested automatically. If you blew through your plan's limit and started consuming usage credits, it drops to five minutes, because now you're paying for the write. On an API key or cloud provider, the default is five minutes and you opt into the hour with ENABLE_PROMPT_CACHING_1H=1.
That explains the slow turn after lunch: the TTL passed, the cache expired, and the next message reprocesses the entire conversation.
When the problem isn't cache, it's compaction
A cache miss costs money and latency. It's annoying, but it's reversible: the next turn is already warm again.
Bad compaction costs something else — decisions. The summary replaces the detailed history as far as the model is concerned. A constraint you agreed on two hours ago, a path already ruled out, the ID of a specific error: if it didn't make it into the summary, it stopped existing for the agent. And then it redoes work you already paid for.
The defense isn't avoiding compaction. It's not letting canonical state live only in the chat.
- Decisions go into an ADR, an issue, or a document in the repository.
- Changes go into a commit or a branch; they don't sit pending until the end of the session.
- Important command results become an artifact in a file.
- Compaction instructions go in the CLAUDE.md: a block saying what to always preserve (modified files, test commands, constraints) survives better than hoping the summary turns out good.
Tool checkpoints, /rewind, and session history are conveniences. Git is the source of truth. It's the same reasoning as persistent state as a mandatory piece of the harness — just applied to you using the tool, not to you building one.
Limitations and things to watch
The cache is narrower than it looks. In Claude Code it is effectively scoped by machine and directory: the system prompt embeds the working directory, platform, shell, OS version, and memory paths. Two sessions in different directories build different prefixes and can't see each other — including worktrees of the same repository. Parallel sessions in the same directory do share.
Agent teams multiply. Each teammate keeps its own context window and runs as a separate instance. The documentation estimates about 7× more tokens than a standard session when the teammates run in plan mode.
The published cost numbers are an average, not a forecast. Anthropic reports about $13 per dev per active day and $150–250 per dev per month in enterprise deployments, with 90% of users under $30 per active day. That varies with codebase, model, and how many instances you run in parallel. If you want the math at today's prices, the table is in Did Claude Code get 5x more expensive? The real price in 2026.
And most important: don't confuse measuring with optimizing. Lowering effort and turning off thinking saves tokens and can cost quality — which ends up more expensive if the agent gets it wrong and you have to redo the work. Optimize what's free first: stable prefix, clean context, filtered output.
Quick FAQ
Is switching models mid-session to save money worth it?
Rarely. The switch invalidates the entire cache and the next request rereads the whole conversation with zero hits. On a large history, the reprocessing can cost more than what the cheaper model saves. If you're going to switch, switch together with a /clear.
/compact or /clear?
/clear when you've changed subjects — it sends no request, costs zero, and is the most effective lever for both quality and cost. /compact only when you need the continuity, and preferably at a break between tasks, not in the middle of one.
Why is the first turn after a break so slow? Expired cache. The TTL is one hour on a subscription, five minutes on an API key or when you're on usage credits. Once it's past, the next request reprocesses the entire conversation as uncached input.
Does having a lot of MCPs really hurt? It depends on whether they're deferred. With tool search active, only the names go into the prefix and a server connecting or disconnecting doesn't take down the cache. Without it, every change to the tool list invalidates everything — and a server can connect or disconnect on its own, through a timeout or reconnection. Cursor's number (−46.9% tokens with dynamic discovery) gives you the scale of the waste when the whole schema sits loaded for nothing.
Conclusion
Saving tokens in a coding assistant isn't about writing less. It's three things: don't touch the prefix in the middle of a task, don't let anything into context that won't be used, and don't keep important state only in the chat.
The rest follows. Pinned model per phase, lean CLAUDE.md, hook-filtered output, /clear between tasks, /rewind instead of /compact when the path has been abandoned.
And there's one metric that sums up the whole setup: the ratio between cache_read_input_tokens and cache_creation_input_tokens. If reads dominate, you're in good shape. If writes won't come down, something in your prefix keeps changing — and that's where your usage limit is going.
Put it in your statusline today. It's the difference between thinking the tool got expensive and knowing why.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã