#Tokens
A decision catalog with numbers, not adjectives: ten tasks that run on a frontier LLM today and fit Jev's three primitives. One million calls in each case costs $46,200 on the LLM and $523 on Jev. It has the payload for each case, the cascade pattern, how to measure the migration in shadow mode without labeling anything, and the four tasks where Jev is the wrong choice.
Your prototype cost R$ 12 on day one. In production, with 200 users, it costs almost R$ 4,000 a month. We open up the token bill of a real agent layer by layer (system prompt, tool definitions, RAG chunks, history, tool results) and show where almost 80% of the money disappears. In the end, prompt compression, history pruning, and context selection cut the bill by 73%, with the before-and-after spreadsheet.
The operational fundamentals of LLMs with no parrot analogy: what the model does at each token, why it's stateless, how the context window degrades long before the limit, and why temperature 0 isn't deterministic. Each concept closes with the architectural consequence it forces you to build, with the official Anthropic and OpenAI docs, the Lost in the Middle paper, and Chroma's context rot study in hand.
You send a one-line question and /usage reports a whole day's worth of consumption. Saving tokens in a coding assistant has nothing to do with prompt size: it's about prefix caching. How it works in Claude Code, Codex, and Cursor, the seven actions that invalidate it without you noticing, how to measure it with cache_read vs cache_creation, and eight levers to stretch the session.