~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / prompt-engineering-is-over-context-engineering $
News

Prompt engineering is over. What came next is called context engineering

LS Lucas Souza · · 13 min read
Prompt engineering is over. What came next is called context engineering

Nobody makes money writing pretty prompts anymore.

The money goes to whoever decides what gets into the context window — and what stays out.

That doesn't mean prompt engineering is trash. It means something more uncomfortable: it stopped being the discipline and became a piece of it. Half the techniques you memorized in 2023 disappeared because the model learned to do them on its own. The other half disappeared because they became API parameters. What's left is small, boring, and exactly what nobody teaches.

In this post you'll see what survived, what turned into folklore, and why in a real system the bottleneck moved: it's no longer the sentence you write, it's the set of tokens that reaches the model on every call.

TL;DR

  • What died: prefill, budget_tokens, "let's think step by step", theatrical personas, threatening CAPS LOCK.
  • What survived: specific instructions, examples, structural delimiting, decomposition, explicit output criteria.
  • What moved layers: reasoning became a dial (effort), format became a schema (output_config.format), memory became infrastructure.
  • Where the real problem lives: order, cutoff, and recency of what goes into the window.
  • Core sources: The Prompt Report (58 techniques cataloged), Context Rot / Chroma (18 models, all of them degrade), Anthropic — context engineering.

If you landed here looking for the basics — what prompt engineering is and how to write instructions, few-shot examples, and output formats that hold up in production — start with the honest guide to prompt engineering. This post is the sequel: what's left of that discipline once the system grows.

The prompt techniques that survived (and the ones that turned into folklore)

The most complete survey on the subject is The Prompt Report: 31 researchers, a systematic review of more than 1,500 papers, a taxonomy with 58 text prompting techniques, 40 for other modalities, and 33 vocabulary terms.

Fifty-eight.

Now answer honestly: how many do you use in production? Probably five. Maybe four.

The ones that survived are the dullest:

  • Specific instructions with acceptance criteria. Not "write a good summary". Instead: "summarize in up to 3 bullets, each with at most 20 words, citing the clause number".
  • Examples (few-shot). Still the cheapest way to transfer format and nuance. With one new catch: examples take up window, and window now has a price.
  • Structural delimiting. In Claude, XML tags (<contexto>, <tarefa>, <formato>) separate things better than loose markdown — the official documentation has recommended this since forever and it still holds.
  • Decomposition. Breaking a big task into three small calls still beats one giant call. Not because the model is dumb, but because compound errors are easier to isolate.

The ones that turned into folklore are precisely the ones that got famous:

  • "Let's think step by step". Chain-of-thought stopped being a text trick and became a button. In Claude you turn on thinking: {type: "adaptive"} and tune the depth with output_config.effort, which goes from low to max (docs). In GPT-5 it's reasoning_effort. You don't ask the model to think anymore. You budget how much it thinks.
  • budget_tokens. The old fixed cap on reasoning tokens is deprecated in Claude Opus 4.6 / Sonnet 4.6 and is rejected with HTTP 400 on Opus 5, Sonnet 5, Opus 4.7, and 4.8. Writing that today isn't a bad prompt: it's a broken request.
  • Prefill. That technique of starting the assistant's response to force the format ({"role": "assistant", "content": "{"}) also returns 400 across the whole 4.6+ family and on the 5 models. What replaced it is real structured output: output_config.format with JSON Schema, and strict: true in the tool definition.
  • "You are a senior expert with 20 years of experience." A theatrical persona doesn't improve accuracy. What does is describing the task, the domain, and the expected format.
  • Threats and CAPS LOCK. "CRITICAL!", "NEVER EVER", "this is VERY IMPORTANT". Anthropic's guidance is blunt: aggressive language produces worse results than calm, specific instructions. The model isn't afraid of you.

There's a pattern here, and it's the point of the post: every prompt technique that was a text trick became an API parameter or became part of the model's training. In practice, what you used to beg for in text you now declare:

{
  "model": "claude-opus-5",
  "thinking": { "type": "adaptive" },
  "output_config": {
    "effort": "high",
    "format": {
      "type": "json_schema",
      "schema": { "type": "object", "properties": { "resumo": { "type": "string" } } }
    }
  },
  "messages": [{ "role": "user", "content": "..." }]
}

Reasoning, depth, and format: three fields. No magic sentence. What's left for the prompt is only what the API has no way of knowing — your task, your domain, your acceptance criteria.

And that's where the conversation changes subject. A prompt solves the single response; what it doesn't solve is the system around it — tool calling, routing, memory, grounding, tracing, evals, and cost. That's exactly the layer we're going to build live, with running code, at AI Engineering Lab 3rd Edition, on September 19 and 20, from 9am to 1pm.

Why a good prompt doesn't save bad context

Now the data point that breaks the illusion.

In July 2025, Chroma published the Context Rot report, by Kelly Hong, Anton Troynikov, and Jeff Huber. They tested 18 frontier models — GPT-4.1, Claude 4, Gemini 2.5, Qwen3 — varying only the input size. The result: all of them degrade as context grows. No exceptions. On 1M-token models, the effect becomes clearly observable between 300,000 and 400,000 tokens.

It's not a bug. It's architecture. As Anthropic explains, the transformer builds pairwise relationships — n² for n tokens. There's a finite "attention budget". Doubling the window doesn't double the attention; it dilutes it.

Translating that into a pain you've already felt:

You have a support agent. The system prompt is flawless. Clear instructions, good examples, defined format, calibrated tone. And it answers wrong.

It doesn't answer wrong for lack of instructions. It answers wrong because, along with your perfect prompt, went 40 messages of history and 12 poorly ranked RAG chunks — and three of those chunks talk about the wrong product in nearly identical vocabulary. The distractor is semantically closer to the question than the right answer is.

No sentence you write in the system prompt fixes that.

Anthropic defines the discipline that does fix it as "the set of strategies for curating and maintaining the optimal set of tokens during LLM inference". Notice the word: curating. Not writing. Curating. It's selection work, not writing work.

What actually goes into the window: order, cutoff, recency

Every call is a budget. You have three decisions — and they're independent.

Order. Position matters. Whatever sits in the middle of a giant block competes with everything around it; whatever opens and whatever closes the context carries disproportionate weight. An instruction at the end of 200,000 tokens of history isn't an instruction, it's a suggestion.

Cutoff. What you leave out is worth more than what you put in. Five well-ranked chunks beat twenty chunks "just in case" — the extra fifteen aren't neutral, they're noise fighting the signal for attention.

Recency. Conversation history ages. A decision made on turn 3 can still hold on turn 40; the tool output from turn 3 almost never does. That's why the APIs got native mechanisms for this: context editing clears out old tool results, compaction summarizes the history, preserving decisions and discarding redundancy.

Each of those three decisions has its own criteria, has trade-offs, and deserves a full post. I wrote one just about that: Context engineering: what goes in the prompt (and what does NOT) — that's where you'll find the breakdown of every type of content competing for the window.

And when the history blows up anyway, the symptom is the classic "the AI forgot what I said". That has a mechanical explanation and four concrete ways out, which I broke down in You hit ChatGPT's token limit: why the AI forgets.

Here the point is a different one, and it's the point of this hub: order, cutoff, and recency aren't prompt. They're architecture.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

System prompt vs. turn instruction

There's a mistake that shows up in almost every project I review: cramming everything into the system prompt. Business rules, customer exceptions, date format, the user's name, today's date. All in there.

And the opposite mistake, less common and just as expensive: repeating the same rule on every turn, burning tokens on something that never changes.

The rule of thumb is simple.

The system prompt is a stable contract. Role, scope, limits, output format, what to do when it doesn't know. Anthropic calls this finding the "right altitude": neither brittle hardcoded logic (which breaks on the first unforeseen case), nor too vague (which guides nothing). Strong heuristics, not a decision tree.

The turn instruction is what changes. The question, the data retrieved just now, the current state of the task.

This isn't nitpicking — it's money. Prompt caching works by prefix, and the rendering order is tools → system → messages. One changed byte in the prefix invalidates everything after it. If you inject datetime.now() into the system prompt, you pay full input price on every request and don't even notice. The way to find out is to look at usage.cache_read_input_tokens: if it's zero on repeated requests, there's a silent invalidator in your prefix.

A new detail that solves a specific case: on the Opus 5 and Opus 4.8 models you can send a system message in the middle of the conversation, as {"role": "system"} inside the messages array, instead of editing the top-level system field. New operator instruction, cached prefix intact.

From prompt to architecture: what changes when prompt engineering becomes a product

In a demo, the prompt is the product. You tweak the sentence, run it again, the output gets better, you post a screenshot on Twitter.

In production, the prompt is a config file surrounded by more expensive things. It still matters. It's just no longer where you spend your time.

What takes its place:

  • Retrieval with evaluated ranking. Having RAG isn't enough. You have to know how many chunks, with which ranking, and measure whether the right chunk showed up.
  • Routing. Not every question deserves the expensive model on high effort. Classifying intent first is the most obvious cost cut that almost nobody makes.
  • Memory outside the window. Anthropic itself describes the pattern: the agent writes structured notes persisted outside the context and rereads only what it needs. Persistent memory with minimal overhead.
  • Evaluation. Without a set of cases and a metric, "the prompt got better" is an opinion. With 50 versioned cases, it's a measurement. If you want to see this with numbers attached, there are three experiments on the same task measured in tokens and dollars in Context engineering beats prompt engineering.
  • Observability and cost. Tracing per request, tokens per route, cost per completed task — not per call.

That's tooling, and tooling has a name, an API, and a price. I've already mapped out what I actually use in Context engineering tools I use in production, with numbers and what goes wrong with each one.

The turning point is this: when the system has more than one turn, more than one source, and more than one tool, the question stops being "how do I write this better" and becomes "what should be in this window right now, and why".

FAQ

So I stop studying prompt engineering? No. You stop studying tricks and start studying specification. Specific instructions, output criteria, well-chosen examples — that's still what separates usable output from mediocre output. What died was the collection of hacks, not the skill of writing a task clearly.

Doesn't a 1M-token window solve everything? No. The Chroma report tested exactly that: 18 models, all degrading as the input grows, with a clear effect on 1M models around 300–400 thousand tokens. A big window is a budget, not a license to spend.

Is it worth using a prompt framework (DSPy, templates, prompt manager)? It is once you already have evals. A framework optimizes what you can measure. Without a set of cases and a metric, it just automates your intuition — faster, and just as baseless.

How do I know whether my problem is prompt or context? Cheap test: take the case that failed, build the call by hand with only the system prompt and the minimum correct context, and run it. If it gets it right, your prompt is fine and the problem is what your pipeline is shoving into the window. If it gets it wrong, then yes, it's the writing.

Does prompt engineering in Portuguese work the same? It does, with one caveat: instructions in Portuguese with data in English (or vice versa) tend to produce mixed output. Pin the output language in the criteria, not in hope.

Conclusion

Prompt engineering didn't end because it failed. It ended because it got absorbed.

The techniques that were worth something became API parameters — effort, output_config.format, strict, adaptive thinking. The ones that weren't became thread folklore. And the hard problem dropped down a layer: it's not the sentence, it's the set of tokens that reaches the model on every call, in what order, with what cutoff, and at what age.

This isn't hype. It's engineering. And it's the part that remains human work, because no API decides for you what's relevant in your domain.

If you want to keep going from here, the natural next step is understanding the selection criteria in detail: Context engineering: what goes in the prompt (and what does NOT).

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing