~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / ai-agent-cost-production-token-bill $
Tutorials

What an AI Agent Costs in Production: The Real Token Bill

LS Lucas Souza · · 13 min read
What an AI Agent Costs in Production: The Real Token Bill

Your prototype cost R$ 12 on day one. In production, with 200 users, it costs almost R$ 4,000 a month. And nobody on the team can explain why.

What an AI agent costs isn't a pricing-page question. The answer is in the spreadsheet nobody opens: how many tokens each layer of the agent (system prompt, tool definitions, RAG chunks, history, tool results) shoves into every single call. You don't pay per question. You pay for context that gets resent, turn after turn.

In this post I open up the token bill for one agent in production, line by line, at today's public prices. Then I pull three levers (prompt compression, history pruning, and context selection) and show the same spreadsheet with the after numbers. Final cut: 73%.

TL;DR

  • What it is: the token cost anatomy of a support agent in production, with a before-and-after spreadsheet.
  • Stack/Models: Claude Sonnet 5 ($2/MTok input, $10/MTok output), tool calling, RAG, prompt caching, context editing.
  • Scenario: 200 users, 18 conversations/month each, 6 turns per conversation, 12 declared tools.
  • Result: from $728.64/month (R$ 3,935) down to $196.67/month (R$ 1,062), without switching models and without cutting a feature.
  • Exchange rate used: R$ 5.40 per dollar. Adjust for today's rate.

Where the tokens go: anatomy of an expensive call

Take turn 6 of a support conversation. The user typed a 120-token sentence. That's all they did. Now look at what your application sent along with it:

Layer Input tokens % of the call
System prompt (persona, rules, format) 1,800 8.5%
Definitions for the 12 tools (JSON Schema) 4,200 19.9%
RAG chunks (5 × 600) 3,000 14.2%
History from the 5 previous turns 2,500 11.8%
Accumulated tool results (4 calls) 9,500 45.0%
User message 120 0.6%
Total input 21,120 100%

The user's question is 0.6% of the bill.

And three layers almost nobody reviews (tool results, tool definitions, and RAG chunks) add up to 79.1%. That's where the money disappears. Not in the model, not in the output, not in "AI is expensive." It disappears in a payload you assembled once, decided was reasonable, and never measured again.

Cost of that single call on Claude Sonnet 5, at $2 per million input tokens and $10 per million output tokens, with a 350-token response:

entrada:  21.120 × $2  / 1.000.000 = $0,042240
saída:       350 × $10 / 1.000.000 = $0,003500
                                     ---------
total do turno 6:                    $0,045740

Four and a half cents. Looks like nothing. Multiply by 6 turns, by 18 conversations, by 200 users, by 12 months, and you have a line item finance is going to ask about.

This math (and the architecture decisions that change it: tool routing, memory, grounding, tracing) is module 11 of the AI Engineering Lab 3rd Edition, a live immersion on September 19 and 20, from 9 a.m. to 1 p.m. That's where we tear down the whole agent, not just the cost layer.

How to measure this without guessing

Don't estimate with tiktoken or with "divide by 4." Measure with the token counting endpoint, layer by layer, by difference:

import anthropic

client = anthropic.Anthropic()
MODELO = "claude-sonnet-5"
PING = [{"role": "user", "content": "."}]

def tokens(**kwargs) -> int:
    return client.messages.count_tokens(
        model=MODELO, messages=PING, **kwargs
    ).input_tokens

piso       = tokens()
com_system = tokens(system=SYSTEM_PROMPT)
com_tools  = tokens(system=SYSTEM_PROMPT, tools=TOOLS)

print("system:", com_system - piso)      # 1.800
print("tools :", com_tools - com_system) # 4.200

Run this in CI. If someone adds a thirteenth tool, the diff shows +380 tokens on every call in the product, forever. That's an architecture decision, not a routine commit.

The hidden cost of history nobody prunes

Here's the part that breaks the mental model of anyone coming from REST APIs: in a conversation, you pay for turn 1 six times.

Context is stateless. Every call resends everything. The bill grows per turn:

Turn Input (tokens) Cumulative
1 9,120 9,120
2 11,520 20,640
3 13,920 34,560
4 16,320 50,880
5 18,720 69,600
6 21,120 90,720

A 6-turn conversation burns 90,720 input tokens for 2,100 output tokens. A 43:1 ratio.

entrada: 90.720 × $2  / 1.000.000 = $0,181440
saída:    2.100 × $10 / 1.000.000 = $0,021000
                                    ---------
por conversa:                       $0,202440

Now the month: 200 users × 18 conversations = 3,600 conversations.

3.600 × $0,202440 = $728,64/mês  →  R$ 3.934,66

Versus R$ 12 on test day (75 short conversations, 3 turns each, with no fat tool results piling up). The prototype wasn't wrong. It was just measuring something else.

The first lever is free: prompt caching

The fixed base (system + tools) is 6,000 identical tokens on every call. Prompt caching bills cache writes at 1.25x the input price (5-minute TTL) or 2x (1-hour TTL), and reads at 0.1x. In other words: from the second call on, that block costs 10% of what it used to.

The detail that trips people up: the minimum cacheable length is 1,024 tokens on Sonnet 5 and 512 on Opus 5. A prefix shorter than that doesn't cache, and the API doesn't complain. It just stays silently expensive.

Two breakpoints cover the entire conversation. A static one at the end of the system prompt (it freezes tools + system, which is the render order) and a moving one at the end of the history:

response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    tools=TOOLS,
    system=[{
        "type": "text",
        "text": SYSTEM_PROMPT,
        "cache_control": {"type": "ephemeral"},   # congela tools + system
    }],
    messages=historico_com_breakpoint_movel + [nova_pergunta],
)

u = response.usage
total = u.input_tokens + u.cache_creation_input_tokens + u.cache_read_input_tokens
print(f"read={u.cache_read_input_tokens} write={u.cache_creation_input_tokens}")

If cache_read_input_tokens comes back zero on repeated requests, there's a silent invalidator in your prefix: a datetime.now() in the system prompt, a tools JSON serialized without stable ordering, a user_id interpolated before the breakpoint. Any byte that changes in the prefix kills everything after it. It's the same mechanism that makes your coding assistant session evaporate, and I've already covered the seven ways to break it in how to save tokens in Claude Code, Codex, and Cursor.

With incremental caching alone, the same conversation drops from $0.2024 to $0.0877. The month goes to $315.72 (R$ 1,704). A 57% cut without touching a single line of the prompt.

Tool results: the biggest context villain

45% of the most expensive call is tool results. It's worth understanding why this happens to almost every agent.

A tool returns whatever the API returns. You call consultar_pedido and get back 2,400 tokens of JSON: full address, status history, carrier data, audit fields, three levels of nested objects nobody reads. The model needed two things: status and estimated delivery date.

Worse: that JSON stays in the history. Turn 3 called the tool. Turn 6 is still paying for it. Four accumulated tool results become 9,500 tokens you resend until the conversation ends.

There are two ways to prune this, and they solve different problems.

In the application, before you send. You control the format. Design the tool's return for the model, not for the log:

def podar_tool_results(messages, manter=3, resumo=200):
    """Mantém os N tool results mais recentes na íntegra.
    Os antigos viram uma linha de sumário."""
    indices = [i for i, m in enumerate(messages) if eh_tool_result(m)]
    for i in indices[:-manter]:
        messages[i] = resumir(messages[i], limite=resumo)
    return messages

On the server, with context editing. The clear_tool_uses_20250919 strategy clears the oldest tool results before the prompt reaches the model, keeping the N most recent:

resp = client.beta.messages.create(
    model="claude-sonnet-5",
    max_tokens=4096,
    betas=["context-management-2025-06-27"],
    context_management={"edits": [{
        "type": "clear_tool_uses_20250919",
        "trigger": {"type": "input_tokens", "value": 30000},
        "keep": {"type": "tool_uses", "value": 3},
        "clear_at_least": {"type": "input_tokens", "value": 5000},
        "exclude_tools": ["consultar_pedido"],
    }]},
    tools=TOOLS,
    messages=messages,
)
print(resp.context_management.applied_edits)
# [{'cleared_tool_uses': 8, 'cleared_input_tokens': 50000}]

Watch the interaction: every clear invalidates the cached prefix. If you clear 800 tokens and rewrite 12,000 tokens of cache, you lost money. That's exactly what clear_at_least is for: it only applies the edit if it's worth the cost of the rewrite.

And the case isn't even purely financial. Anthropic's documentation is blunt: "context must be treated as a finite resource with diminishing marginal returns." There's the context rot effect: as the token count grows, the model's ability to retrieve information from that context drops. You pay more for the agent to get more things wrong.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

Compression and pruning in practice, with before-and-after numbers

Three levers, applied to the same turn 6 call.

1. Prompt compression. The 1,800-token system prompt had four few-shot examples doing the same thing. Now there are two: 1,650. All 12 tools were declared every time, on every call, even when the conversation was only about tracking. Domain-based routing loads 5: from 4,200 down to 1,750.

2. History and tool result pruning. The 3 most recent tool results stay intact, the old ones become a one-line summary: from 9,500 down to 4,100. Turns 1 through 3 of the history become a one-paragraph summary: from 2,500 down to 1,400.

3. Context selection. The 5 RAG chunks went into every call by default. With reranking and a score cutoff, 2 go in, and only when the query calls for retrieval: from 3,000 down to 1,200.

Turn 6, after:

Layer Before After Delta
System prompt 1,800 1,650 -8%
Tool definitions 4,200 1,750 -58%
RAG chunks 3,000 1,200 -60%
History 2,500 1,400 -44%
Tool results 9,500 4,100 -57%
User message 120 120 -
Total 21,120 10,620 -49.7%

The whole conversation drops from 90,720 to 46,020 input tokens. With the two cache breakpoints on top:

Scenario Per conversation Per month (3,600) In reais
Baseline, no cache, no pruning $0.2024 $728.64 R$ 3,934.66
Prompt caching only $0.0877 $315.72 R$ 1,704.89
Caching + compression + pruning $0.0546 $196.67 R$ 1,062.02

A 73% cut. Same model, same feature, same response quality.

And notice what changed shape: in the optimized scenario, the $0.021 of output became 38% of the conversation's cost. While the input was obese, the output was statistical noise. After pruning, it becomes the next bottleneck. That's where verbosity limits, a lower output_config.effort on simple routes, and structured output to cut the preamble come in.

If your use case doesn't need a real-time response (batch classification, overnight enrichment, report generation), the Batch API cuts another 50% on top of that, with most batches finishing in under an hour.

The margin math: when the product doesn't pencil out

Now the part that decides whether the product exists.

A R$ 79 subscription per user. 200 users. MRR of R$ 15,800.

Token cost % of revenue Cost per user
Baseline R$ 3,934.66 24.9% R$ 19.67
Optimized R$ 1,062.02 6.7% R$ 5.31

SaaS lives on gross margins of 70% to 80%. With inference eating 25% of revenue, before infra, observability, support, and payroll, there's no product left. At 6.7%, there is.

But the average lies. What kills margin is the tail.

Profile Conversations/month Optimized cost Revenue Outcome
Median 18 R$ 5.31 R$ 79 healthy
P95 90 R$ 26.55 R$ 79 tight
P99 300 R$ 88.50 R$ 79 R$ 9.50 loss

At baseline, that same P99 user cost R$ 327.89 a month to pay R$ 79. A single power user wiped out the margin of four median customers.

Three consequences, and they're product consequences, not engineering ones:

  1. Flat unlimited pricing only works if the cost ceiling is architectural. Either you cap turns per conversation, or you cap conversations per cycle, or you charge by usage. There is no fourth option.
  2. Cost per user needs to be on a dashboard, not in a quarterly report. Log input_tokens, cache_read_input_tokens, cache_creation_input_tokens, and output_tokens per request, with user_id and route. One token_ledger table and one chart. It's half a day of work and it pays for the year.
  3. Tokens are just one of the buckets. Runtime evals, observability, vector store, queue, people. If you want the full bill for an agent over six months, that's what I laid out in What an agent costs in production in 2026: a real TCO spreadsheet. This post is just its inference line, broken out into layers.

Quick FAQ

So how much does an AI agent cost in production?

It depends almost entirely on how many tokens you resend per turn, not on the model's price. In this post's scenario, the same agent costs between R$ 1,062 and R$ 3,935 a month for the same 200 users. The variable isn't the vendor. It's your context architecture.

Isn't the price per token dropping? Isn't it better to wait?

The price per token drops, but consumption per task rises faster, because newer agents take more turns and use more tools. That's a market discussion I broke down in Is AI going to get more expensive?. For next month's spreadsheet, list price is the variable you don't control. Context is the one you do.

Doesn't switching models fix this faster than optimizing context?

Sometimes, and it's worth measuring. But cache is scoped per model: a two-model cascade opens two cache namespaces and you lose reuse in both. Do the context hygiene first (it's free), measure again, and only then switch models with data in hand.

Why is my cache_read_input_tokens always zero?

Something in the prefix changes on every call. The usual suspects: a timestamp in the system prompt, non-deterministic ordering in the tools JSON, a session ID interpolated before the breakpoint, or a prefix shorter than the model's minimum cacheable length (1,024 tokens on Sonnet 5). Log the raw usage from two back-to-back calls and compare byte by byte.

Closing the spreadsheet: what an AI agent costs in your case

The bill for an agent in production isn't a number you discover. It's a number you design.

Every tool that gets added, every few-shot example someone pastes into the system prompt, every extra RAG chunk in the top-k: each of those is a recurring cost decision, multiplied by every turn of every conversation of every user. And almost none of it shows up in code review.

The path is always the same: measure per layer, cache what's stable, prune what's old, select what's relevant. In that order. In this scenario it came out to a 73% cut without switching models.

Once the inference line is closed, the next step is to look at the five buckets that are left (evals, observability, infra, data, and people) and decide build vs. buy with numbers in hand.

Do you know how much the last call your agent made cost? If the answer is "no," that's the first bug to fix.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing