Fine-Tuning, RAG, or Prompt: Which One Your Problem Actually Calls For
"Let's train our own model."
I've heard that line in architecture meetings more times than I can count. And the next question almost always kills the idea: tell me what you're trying to teach that good context wouldn't solve for 3% of the price.
I'm not being difficult. It's that fine tuning has become the wrong reflex. When the model gets something wrong, a lot of teams instinctively reach for training. But in most cases the problem isn't in the model's weights, it's in what you put (or failed to put) inside the context window.
This post is a decision framework, not an explanation of concepts. You'll walk away with four objective questions for choosing between prompt, RAG, and fine-tuning, the real cost of each path with numbers, the three rare cases where training actually wins, and the one-line diagnostic that separates a knowledge problem from a retrieval problem.
TL;DR
- The right order is: prompt → RAG → evals → fine-tuning. In that sequence, and you only move on when the previous step has proven it hit the ceiling.
- The RAG vs fine tuning fight is badly framed: fine-tuning changes HOW the model answers, RAG changes WHAT it knows. Confusing the two is the mistake that gets expensive.
- Real cost: prompt with caching cuts the inference bill by ~62% and charges nothing for setup. Fine-tuning costs thousands of dollars before the first useful answer and then makes inference up to 50% more expensive forever.
- The news that changes the math in 2026: OpenAI is shutting down its fine-tuning platform. Active customers stop creating jobs on January 6, 2027.
- Scientific evidence: training in new knowledge increases hallucination linearly. That's not an opinion, it's an EMNLP paper.
What each one actually changes in behavior
The three act on different layers of the system. That distinction settles half the decisions on its own.
Prompt changes the instruction. You reconfigure behavior on every request. Marginal cost of a change: zero. Iteration time: seconds. Reversible with a git revert.
RAG changes the context. You inject facts at question time and the model stays exactly the same. If you're still early on this part, I've already covered what RAG is and where it ends and the chunking, embeddings, and search pipeline in dedicated posts. Here I'll treat RAG as a black box and talk only about decision and cost.
Fine-tuning changes the weights. You alter the probability distribution of the next token. It's permanent, expensive to reverse, and requires a new training cycle for every behavior tweak.
Now the sentence that's worth the whole post:
Prompt and RAG operate at runtime. Fine-tuning operates at build time.
Any information that changes after deploy has to come in at runtime. If you trained the model on August's price table, in September you have a model that lies about pricing with conviction. No prompt fixes that, because the error is in the weights.
This is literally the first module of the AI Engineering Lab 3rd Edition: before tool calling, memory, or tracing, deciding which layer your system's behavior is configured in. It's two live days, September 19 and 20, from 9am to 1pm, on Meet, building production agent architecture.
The decision tree in 4 questions
If all you want to know is when to use fine tuning, the honest answer is in these four questions. Run them in order: the first one that settles it, settles it.
1. Does the missing data change over time?
If the correct answer today is different from the one three months ago, fine-tuning is out. There's no debate here. Price, inventory, internal policy, API documentation, order status: all of that is runtime. Go with RAG or a tool call.
Fine-tuning freezes things. And frozen data in production isn't an outdated model, it's a confident and wrong model, which is a lot worse.
2. Is the model getting the FACT wrong or the FORM wrong?
This is the question most people skip.
- Gets the fact wrong (makes up a number, cites a clause that doesn't exist, attributes a feature to the wrong product) → retrieval problem. RAG.
- Gets the form wrong (returns broken JSON, ignores the tone instruction, writes three paragraphs when you asked for one line, picks the wrong route) → behavior problem. Prompt first. Fine-tuning only if the prompt proves it can't get there.
OpenAI says the same thing in its official optimization guide: fine-tuning is for classification, format, tone, and instruction-following failures. None of those is knowledge.
3. Do you have an eval?
If you don't have a scored evaluation dataset, you have no way of knowing whether training improved anything. You'll be swapping one gut feeling for another, with a four-figure cost in between.
Fine-tuning without an eval isn't engineering, it's faith. And the test is an honest one: if you can't say "today I'm at 71% accuracy on my 120 cases," you're not ready to train. You're ready to build the eval. I've already shown how to set up that pipeline in CI.
4. Have you really exhausted the prompt?
In practice, almost nobody has. Before training, have you already:
- added 5 to 10 real few-shots, pulled from cases that failed?
- forced
structured output/ JSON schema instead of asking for a format in plain text? - split the task into two calls instead of one heroic call?
- tested the bigger model for a week to see where the ceiling is?
- fixed the retrieval, instead of blaming the generator? (the real causes are here)
If any item is left unchecked, you haven't finished the cheap path.
What each path costs, with numbers
Let's anchor this in a concrete scenario: a technical support agent, 100,000 requests per month, a system prompt with instructions and few-shots adding up to 6,000 fixed tokens, a user question of 200 tokens, a response of 400 tokens. Claude Sonnet 4.6 pricing ($3 per 1M input, $15 per 1M output).
Path 1: prompt
Without caching, the 6,000 fixed tokens are billed 100,000 times:
input: 6.200 tok × 100.000 = 620M × $3/M = $1.860
output: 400 tok × 100.000 = 40M × $15/M = $600
total = $2.460/mês
With prompt caching, a cache read costs 0.1x the input price and a 5-minute cache write costs 1.25x. Assuming a 95% hit rate on the fixed block:
cache read: 6.000 × 95.000 = 570M × $0,30/M = $171,00
cache write: 6.000 × 5.000 = 30M × $3,75/M = $112,50
input novo: 200 × 100.000 = 20M × $3,00/M = $60,00
output: 40M × $15,00/M = $600,00
total = $943,50/mês
62% savings from a three-line code change. Setup: zero new infra, days of engineering. This is the number most teams never go after before proposing to train a model.
Path 2: RAG
The cost of embeddings scares a lot of people, and it's the cheapest part of the bill. 50,000 documents of 800 tokens come to 40M tokens: with text-embedding-3-small at $0.02 per 1M, that's $0.80, once. The month's queries cost pennies.
The real cost of RAG is somewhere else:
- Vector infra (Postgres with pgvector, dedicated instance): $50 to $200/month.
- Managed reranker, at ~$2 per 1,000 searches: $200/month for the 100,000 searches.
- Retrieved context tokens, which is the line item that hurts: 3,000 tokens per request that change with every question and therefore don't cache. That's 300M tokens at $3/M = $900/month.
An incremental ~$1,150/month on top of path 1. Engineering: 2 to 6 weeks until it becomes something reliable in production.
Path 3: fine-tuning
Here the math changes in nature, because the dominant cost isn't API cost.
Training, using OpenAI's price of $25 per 1M training tokens on gpt-4.1: a dataset of 5,000 examples at 1,200 tokens each is 6M per epoch, 18M across three epochs. $450 per run. And nobody gets it right the first time: 5 to 10 runs to reach an acceptable checkpoint, meaning $2,250 to $4,500.
Dataset curation, which is the cost nobody puts on the slide: 5,000 examples reviewed by people who know the domain, at 60 examples per hour, comes to 83 hours of work. At R$ 150/h that's R$ 12,450. It's the most expensive item on the entire list, and it doesn't show up in any token calculator.
More expensive inference forever: fine-tuned gpt-4o costs $3.75 for input versus $2.50 for the base model, and $15 for output versus $10. A 50% surcharge on every request, every month, for as long as the model lives.
And the maintenance cost: every behavior change means a new dataset, a new training run, a new validation. It's not a git revert, it's a sprint.
The table
| Setup | Monthly cost | Time to production | Rollback | |
|---|---|---|---|---|
| Prompt + caching | ~$0 | $943 | days | commit |
| RAG | days of engineering | ~$2,100 | 2 to 6 weeks | turn off the tool |
| Fine-tuning | $2,250 to $4,500 + R$ 12,450 in curation | inference +50% | 1 to 3 months | retrain |
Look at the "rollback" column. It's the most important one and the one the fewest people consider when deciding.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãThe 3 rare cases where fine tuning wins
It's not never. It's rare. And all three cases share the same signature: what you're teaching is behavior, not fact, and the goal is almost always to cut cost, not to raise the quality ceiling.
1. Distillation to run cheap at scale
You have a giant prompt running on a big model and working well. You use that model to generate 20,000 high-quality examples and train a small model to imitate the behavior.
The gain here is cost and latency, not intelligence. You trade $15 of output for $5, and a 4-second response for 800ms. If your volume doesn't justify that delta, this isn't your case.
2. A domain with language the pretraining barely saw
A medical report full of hospital shorthand, a legal filing from a specific niche, an internal DSL at your company that lives in three private repos. The model doesn't have the patterns, and no amount of few-shot makes up for it.
Notice that even here the gain is about how to speak, not what to know. The facts still come from retrieval.
3. Very high-volume classification with stable labels
Content moderation, ticket routing, intent triage. The label set hasn't changed in a year, the volume is in the millions per month, and latency matters.
A small trained model beats a prompt on a big model by an absurd margin on cost per call. This is the most legitimate of the three cases and, as it happens, it's exactly the one OpenAI still points fine-tuning at in its official guide.
The 2026 warning
Before you architect on top of this, look at the calendar. OpenAI is shutting down its fine-tuning platform:
- May 7, 2026: organizations that have never trained can no longer create jobs.
- July 2, 2026: also blocked for anyone who hasn't run inference on a fine-tuned model in the last 60 days.
- January 6, 2027: active customers stop creating new jobs.
Inference on already-trained models continues until the base model is deprecated. OpenAI's own justification was blunt: the new models have become capable enough to make a good chunk of fine-tuning unnecessary.
That doesn't kill fine-tuning (open weights with LoRA are alive and well, and that's where this use case migrates). But it does kill the idea that managed fine-tuning is the safe default choice for the long run.
The classic mistake: training to solve a retrieval problem
This is the mistake that closes the post because it's the most expensive of all: the team wants the model to "know" the company's knowledge base, so they dump the entire documentation into a training dataset.
It doesn't work. And there's literature on it.
The paper Fine-Tuning or Retrieval? (EMNLP 2024) compared the two methods for injecting knowledge. On the current events task, RAG scored 0.875 against 0.504 for fine-tuning on Mistral. More than double, with the base model untouched.
FineTuneBench went deeper and measured the commercial APIs: 37% generalization accuracy for absorbing new information, and 19% for updating knowledge the model already had. Gemini 1.5 Flash and Pro simply didn't learn.
And the most serious one, from Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?: examples with new knowledge are learned more slowly and, as they're learned, they linearly increase the model's tendency to hallucinate. You don't just fail to teach the fact: you degrade the model's ability to use what it already knew.
Translated to production: you spent R$ 15k to end up with a model that gets more things wrong.
The one-line diagnostic
There's a test that settles this question in five minutes, and it's embarrassingly simple. Take the case that failed and paste the right document straight into the prompt, by hand:
Contexto:
<cola aqui, na unha, o trecho do documento que contém a resposta>
Pergunta:
<a pergunta que o modelo errou em produção>
Run it. And read the result like this:
- Got it right with the document pasted in → you have a retrieval problem. The model is capable, it just never got the text. Fix search, chunking, or the reranker. Training here is burning money to fix a
SELECTbug. - Got it wrong even with the document in front of it → now it really is a behavior or capability problem. Now the conversation about a better prompt, a bigger model, or fine-tuning starts to make sense.
In my experience, this test sends the overwhelming majority of "we need to train a model" requests back to the retrieval backlog. It's why the reranker is, almost always, the best return per R$ invested before anything more sophisticated.
Quick FAQ
Can you use RAG and fine-tuning together? You can, and in mature systems it's common: fine-tuning controls format and reasoning, RAG controls the facts. But that's an optimization for people who already have both working separately, not a starting point. I covered the hybrid architecture here.
How many examples do I need for fine-tuning to be worth it? OpenAI's technical minimum is 10, and the recommendation is to start with 50 well-crafted ones. But for consistent behavior change in production, the realistic range is 1,000 to 10,000 curated examples. If you have 200 examples, they belong in the prompt as few-shots, not in a training job.
Does prompt caching work with context retrieved by RAG? Only on the stable part. The fixed block (instructions, schema, few-shots) caches well; the retrieved chunks change with every question and get no benefit from the cache. That's why it pays to order the prompt from most static to most dynamic, with the RAG content at the end.
What if I want to run fine-tuning on an open weight model? Then it's a different game: LoRA on an open model takes API cost out of the equation and gives you control back. The entire decision tree still applies, especially question 1. An open model trained on volatile data is wrong just the same.
Conclusion
The framework fits in one line: prompt until you hit the ceiling, RAG when facts are missing, fine-tuning when behavior is missing and the volume justifies the bill.
And before any of the three, an eval. Without an eval you're not choosing an architecture, you're choosing based on a feeling.
Fine-tuning isn't dead. But in 2026 it stopped being the default answer and became a cost optimization for mature use cases, with a curated dataset and metrics in place. The path that almost always pays off fastest is still the most boring one: fix the retrieval and write a better prompt.
If you want to dig into the selection criteria from another angle (data volatility, traceability, and corpus size), the post on when to use RAG, fine-tuning, or context pairs well with this one.
Now go run the pasted-document test on your failing case. My money's on retrieval.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã