What Is an LLM: The Explanation for People Who Build, Not People Who Post
You don't need to understand transformers to use an LLM. You need to understand three things — and none of them is the transformer.
Most content about what an LLM is stops at the stochastic parrot, the souped-up autocomplete analogy, and an attention diagram nobody can use to make a single decision. That's an explanation for people who are going to post about AI. If you're going to build with AI, what you need to know is something else: what the model mechanically does at each step, what it doesn't do for you, and which part of the bill is left for your system to pay.
This post is the operational foundation. Four concepts — token, state, window, and sampling — and, for each one, the architectural consequence it forces you to accept. No parrot analogy. With the official docs in hand.
What an LLM is, in short (TL;DR)
- What it is: LLM stands for large language model. In practice, a model that, given everything in its input, estimates the probability of the next token and picks one. Then it repeats. That's the entire operation.
- The three fundamentals that matter: token-by-token prediction, a total absence of state between calls, and a context window that is finite and degradable.
- What it is NOT: memory, a database, a deterministic function, or an API with a format guarantee.
- Architectural consequence: state, retrieval, validation, and evaluation are your code's responsibility. The model isn't going to do any of that for you.
- Reference docs: Anthropic's official glossary and the context window guide.
What the model actually does at each token
Start with the no-frills definition. Anthropic itself describes pretraining like this: autoregressive models are trained "to predict the next word, given the previous context of the text in the document" (official glossary).
That's it. At each step, the model takes in the entire sequence that exists so far, produces a probability distribution over the whole vocabulary, picks a token, appends it to the end of the sequence, and runs again. There's no plan stored somewhere. There's no persistent internal draft. Every token is a local decision, conditioned on everything that has already come out.
And a "token" is not a word. A token is the unit the model's tokenizer defined — it can be a word, a subword, a character, or a byte. For Claude, the official docs give you the yardstick: a token represents approximately 3.5 characters in English, "though the exact number can vary depending on the language used."
That "depending on the language" is the part that costs money in Brazil. A study presented at NeurIPS 2023 measured the same text translated into several languages and found tokenization differences of up to 15x between language pairs. Portuguese isn't the worst case, but it pays the tax: in realistic prose with the cl100k_base tokenizer, Portuguese consumes about 1.46x the tokens of the equivalent English, and with o200k_base the gap sits around +38% (reproducible benchmark).
Translated into a real decision: if your system prompt is 3,000 tokens in Portuguese and gets resent on every request, it would cost around 2,000 in English. In a product with volume, that's a line item on a spreadsheet, not trivia.
The architectural consequence: the model generates left to right, one token at a time, without looking ahead. It doesn't know the answer before it starts writing. That's why "ask for JSON in the prompt" was never a guarantee — it's a probability. And that's exactly why structured output and tool calling with a strict schema exist: they constrain the sampling space on the server instead of hoping the distribution cooperates. You don't control the output. You control the input and the constraints.
This fundamental is where the architecture decisions come from that we'll be building live at the AI Engineering Lab 3ª Edição, on September 19 and 20: tool calling, structured output, routing, memory, grounding, tracing, evals, and cost. It's two days from 9 a.m. to 1 p.m., online on Meet, with code on screen and bugs being debugged on the spot.
Why it's stateless (and what that forces you to build)
This is where most people get fooled, because the chat interface lies very well.
OpenAI's docs are literal about it: "each text generation request is independent and stateless" — and to do multi-turn conversation "you include the model's previous response as input, and append that input to the next request" (conversation state guide).
There is no session on the model's side. There is no "it remembered." What there is: your client reassembling the entire conversation and resending all of it, from scratch, on every turn:
# turno 3: você não "continua" nada. Você reenvia a conversa inteira.
messages = [
{"role": "user", "content": "qual o prazo do pedido 4471?"},
{"role": "assistant", "content": "O pedido 4471 sai dia 12."},
{"role": "user", "content": "e o 4472?"},
{"role": "assistant", "content": "O 4472 sai dia 15."},
{"role": "user", "content": "manda os dois pro meu email"}, # <- só isso é novo
]
The model only knows about order 4471 because that text is physically present in request number 3. Remove the line and it has no idea what you're talking about.
Now the math, which is the part nobody shows. In a 20-turn conversation with ~500 tokens per turn, the processing isn't 20 × 500 = 10,000 input tokens. It's the cumulative sum: 500 × (20 × 21 / 2) = 105,000 tokens. Ten and a half times more. The cost of a conversation grows quadratically with the number of turns, not linearly.
That's why prompt caching exists and isn't an optional optimization in production: reusing the already-processed prefix drops the read cost to a fraction of base input. If you've never looked at this math up close, we wrote about how to stop breaking the cache.
The architectural consequence: memory is a feature of your system, not of the model. Every time an agent "remembered" something between runs, someone wrote code that persisted it and rehydrated the context on the next call. There's no magic — there's a write to disk, or a table, or a vector. We already took this apart in an agent left a note for the next one. If your solution design doesn't have an explicit answer to "where does the state live," you don't have a solution design.
Context window: the resource you manage
If the model has no state and you resend everything, there's a ceiling. That ceiling is the context window — described in Anthropic's docs as the model's "working memory."
First, what counts toward it. Everything: the system prompt, every message in the array, tool results, images, documents, the definitions of your tools, and the output the model generates, including thinking tokens. If it's in the request, it takes up space.
The current numbers, from the official docs: Claude Opus 5, Opus 4.8, 4.7, 4.6, Sonnet 5, Sonnet 4.6, and Fable 5 have a 1 million token window, with up to 128K output tokens per request. Earlier models, like Sonnet 4.5 and Haiku 4.5, sit at 200K.
And when it overflows? If the input alone already exceeds the window, you get a 400 invalid_request_error with "prompt is too long". If input + max_tokens exceeds it, models from 4.5 up accept the request and stop generation with stop_reason: "model_context_window_exceeded" — which is a silent error and far more dangerous, because it looks like a response.
Now the point that separates people who read the docs from people who read the marketing number: the technical limit is not the useful limit.
Anthropic itself names the phenomenon in its documentation: "as token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available."
This isn't vendor opinion. There's literature:
- Lost in the Middle (Liu et al., TACL 2024): performance traces a U-shaped curve. The model makes good use of what's at the beginning and the end of the input, and degrades significantly when the relevant information is in the middle — "even for explicitly long-context models."
- Context Rot (Chroma, July 2025): 18 frontier models tested — GPT-4.1, Claude 4, Gemini 2.5, Qwen3 among them — and all of them get worse as the input grows, non-uniformly and long before the window limit (report).
In other words: stuffing 900K tokens into a 1 million window is technically valid and operationally dumb.
The architectural consequence: the window is a budget, not a closet. Every token you put in competes for attention with all the others. That's why RAG and retrieval don't exist to "make it fit" — they exist to put a small amount of the right stuff in the window. And that's why prompt engineering became context engineering: the job stopped being writing the magic sentence and became deciding, on every call, what deserves to take up space.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãTemperature, top-p, and the illusion of control
Every LLM tutorial has the "tune the temperature" section. Almost none of them tell you what it doesn't do.
The basics, straight from the OpenAI API reference: temperature ranges from 0 to 2 and controls how much the probability distribution is flattened before the pick — high values make the output more random, low values more focused. top_p is nucleus sampling: instead of changing the shape of the distribution, it cuts off the tail and only considers the tokens that add up to that probability mass. And the docs' own recommendation: "we generally recommend altering this or top_p, but not both" — because both act on the same distribution and the combined effect is hard to predict.
Up to here, it's what everybody repeats. Now the two things almost nobody talks about.
First: temperature=0 is not deterministic. It's in Anthropic's docs, spelled out: "Even with temperature set to 0, the results will not be fully deterministic and identical inputs can produce different outputs across API calls."
And you can measure it. Thinking Machines Lab published a clean experiment in September 2025: 1,000 completions of the same prompt, on the same model (Qwen3-235B-A22B-Instruct-2507), at temperature 0. Result: 80 unique responses. The thousand completions were identical up to token 102 and diverged at 103 — 992 wrote "Queens, New York" and 8 wrote "New York City". The root cause isn't the model being "creative": it's a lack of batch invariance in the inference kernels. The size of the batch your request landed in on the server changes the order of the floating-point reductions, and that changes the result.
You don't control the batch size. Nobody consuming an API does.
Second: on frontier models, these parameters have ceased to exist. From Anthropic's docs on thinking: on Claude Fable 5, Mythos 5, Opus 5, Opus 4.8, Opus 4.7, and Sonnet 5, non-default values of temperature, top_p, or top_k "return a 400 error on every request, regardless of whether thinking is used or not." The knob half the tutorials teach you to turn simply isn't on the panel anymore for the most capable models.
The architectural consequence: if your system's correctness depends on temperature=0, you don't have a guarantee — you have statistically good luck that's going to fail in production at peak hours, when the server's batch is different. Determinism, when you need it, gets built outside the model: a schema validated in your code, retry with verification, idempotent writes, and an eval running against a set of cases. Not in a sampling parameter.
What this means when you're designing the architecture
Put it all together and the design becomes obvious. Each fundamental turns into an obligation:
| The model is like this | You're forced to build |
|---|---|
| Predicts one token at a time, with no plan up front | Format constraints on the server (structured output, strict schema), never just an instruction in prose |
| Has no state between calls | An explicit memory layer: where the state lives, who writes it, who rehydrates it |
| Has a finite window and degrades long before reaching it | Curated retrieval — a small amount of the right stuff — and not "throw in everything the window can take" |
| Samples probabilistically, even at temperature 0 | Validation, retry, and eval outside the model |
Notice the pattern: all four rows on the right are your code. None of them comes free with the API.
That's why an "AI engineer" isn't someone who writes pretty prompts. It's someone who knows which piece of the problem the model solves and builds the engineering around the piece it doesn't. The model is a probabilistic component inside a system that needs to be deterministic enough for a customer to trust it.
And there's a nice side effect to understanding this: you stop buying hype. When a new model ships with a bigger window, you don't ask "how many tokens?", you ask "where does it start degrading?". When someone sells you an agent with "infinite memory," you ask where the state is persisted. When a framework promises structured output, you ask whether it's a schema on the server or an instruction in the prompt.
Fundamentals aren't the boring topic before the fun topic. Fundamentals are what make you ask the right question.
Quick FAQ
How does an LLM work, without getting into transformers? In a loop. It takes in the entire input sequence, computes a probability distribution over the whole vocabulary, picks a token, appends it to the end of the sequence, and repeats the process from scratch. There's no plan up front and no internal draft kept between one token and the next.
Are LLMs and generative AI the same thing? No. Generative AI is the umbrella (image, audio, video, text). LLMs are the family of language models trained on text to predict the next token. Every LLM is generative AI; not all generative AI is an LLM.
Does the model learn from what I send through the API? Not by default. Training and inference are separate processes: your request doesn't change the model's weights. What looks like "learning" in a conversation is just the information being inside the context window of that request — and it's gone the moment you don't resend it. Changing behavior persistently requires fine-tuning or a memory layer of your own.
Does a 1 million token window mean I can send the whole repository? You can, and you probably shouldn't. The limit is technical, not useful: Lost in the Middle and Context Rot show degradation well before the ceiling, in every model tested. Retrieving the right 20 files usually beats sending all 2,000.
So how do I make the model's output predictable? You don't make the model predictable — you make the system tolerant. Structured output with a strict schema to constrain the format, validation in your code, retry when validation fails, and a set of evals with real cases to catch regressions. Temperature is the last and weakest of the controls, when it exists at all.
The next step
An LLM is a next-token probability estimator, with no state, a finite attention budget, and sampling you don't fully control. Everything else — memory, format, truth, consistency — is engineering that's left to you.
That's the floor. It's on top of it that the things we write about here make sense: retrieval, tool calling, evaluation, cost, routing. If any of those terms still sounds fuzzy, the AI Engineer glossary with 30 terms sorts out the vocabulary fast.
And one provocation to close: in 2026 the model debate became a commodity — everybody has a big window, everybody has tool calling, prices drop every month. What still separates a product from a demo is exactly what this whole post described: what you build around the probabilistic piece.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã