~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / llm-grounding-ai-hallucination $
News

LLM Grounding: The 4 Layers That Pin the AI's Answer to Your Data (and Validate Before Delivery)

LS Lucas Souza · · 13 min read
LLM Grounding: The 4 Layers That Pin the AI's Answer to Your Data (and Validate Before Delivery)

The model didn't lie. It completed.

The difference between those two things is what separates a demo from a product.

You open the ticket, check the log, and there it is: the agent made up an order number that never existed. Confidently. Nicely formatted. In a tone that would sail through any copy review. And the paying customer on the other end believed it.

AI hallucination isn't a model bug. It's a missing tether. And the default reaction — blame the model, switch providers, lower the temperature, ask "please don't make things up" — fixes nothing, because the diagnosis is wrong.

In this post you'll see why the name "hallucination" gets in the way of the diagnosis, which techniques pin the answer to data that's yours, how to validate in real time instead of praying inside a try/catch, and what will still slip through even after all of that.

TL;DR

  • What it is: hallucination is the model completing plausible text when it has nothing to hold on to. Expected statistical behavior, not a defect.
  • What fixes it: grounding (pinning it to your data), structured output with a strict schema, real-time validation, and verified source citation.
  • What doesn't: switching models, lowering temperature, asking "don't make things up" in the prompt.
  • Base source: Why Language Models Hallucinate — Kalai, Nachum, Vempala, and Zhang (OpenAI + Georgia Tech), September 2025.

Read first: this post is the architecture layer. If you still want to understand the cause — why the model makes things up and what to do about it from the product angle — start with why AI hallucinates and how to reduce hallucination in your product. Here we assume you already know it makes things up and you want the system that holds the answer in place.

Why "AI hallucination" is the wrong name for the problem

"Hallucination" suggests delirium. It suggests the model had an episode, saw things that weren't there, went off the rails. The name is bad because it implies there's a "sane" state to go back to.

There isn't.

What actually happens is more boring and far more useful to understand: the model is a completer. It estimates the next token. When the information you asked for isn't in the context and isn't solid in the weights, it has no "missing data" button. It completes with whatever is statistically plausible.

The paper Why Language Models Hallucinate, published in September 2025, puts it bluntly: hallucinations "originate simply as errors in binary classification." If an incorrect statement can't be distinguished from fact during pretraining, natural statistical pressure produces hallucination. It's not a mystery. It's the math working as designed.

And then there's the part that hurts more. The authors show that hallucination persists because we reward guessing:

Language models are optimized to be good test-takers, and guessing when uncertain improves test performance.

The model is a good test-taker. On a multiple-choice exam, guessing has positive expected value and answering "I don't know" is worth zero. The benchmarks that dominate the leaderboards work exactly like that. The model learned to never leave a question blank because we taught it to never leave a question blank. So much so that the mitigation the paper proposes isn't a new hallucination eval: it's changing the scoring of the benchmarks that already exist.

Notice what that changes on your side. If hallucination comes from a missing tether and an incentive to guess, the fix isn't in the model. It's in the architecture around it: what goes into the context, the format you force on the output, and what you verify before showing it to the user. That's systems engineering, not prompt engineering. It's exactly the layer we build live at AI Engineering Lab 3rd Edition, on September 19 and 20, with grounding, tracing, evals, and cost on top of a running system on screen.

Grounding: pinning the answer to data that's yours

Grounding is a simple idea: the model's answer has to depend on a source of truth you control, not on its memory.

The source can be a retrieved document, a tool's return value, a row in your database, the text of an internal policy. The point isn't where it comes from. It's that it exists, that it's in the context, and that the instruction makes it explicit that nothing outside of it counts.

Three tethers handle most cases.

1. Explicit restriction on outside knowledge. Anthropic's documentation on reducing hallucinations is direct: instruct the model to use only the information in the provided documents, not its general knowledge. Sounds obvious. Almost nobody writes it in the prompt.

2. Permission to not know. Same doc, and it's the highest-return technique per line of prompt: explicitly allow the model to admit uncertainty. Remember the paper? It learned that "I don't know" scores zero points. In your system, "I don't know" is worth far more than a wrong answer — but you have to say that somewhere.

Responda usando apenas o conteúdo entre <documentos>.
Se a informação necessária não estiver ali, responda exatamente:
"Não encontrei essa informação nos documentos disponíveis."
Não complete com conhecimento geral.

3. Extracting verbatim quotes before the answer. For long documents, the recommendation is to flip the order: ask for the verbatim quotes first, then the analysis on top of them. You force the model to go through the actual text before reasoning, instead of reasoning and then hunting for support.

And RAG, where does it fit in? As one of the tethers — the one that handles the specific case of "the data is too big for the context and changes often." RAG is the retrieval mechanism. Grounding is the property you want. You can have grounding without RAG (a tool call to an endpoint, a query against the database) and you can have RAG with no grounding at all: it retrieves five chunks, ignores all five, and answers off the top of its head. If that foundation isn't solid yet, what is RAG covers how it works, and your RAG doesn't work? the real causes covers what breaks in practice.

Real-time validation vs. praying in the try/catch

Grounding raises the odds that the answer is right. It doesn't guarantee format, and it doesn't guarantee the content matches the source.

This is where most production systems have a try/catch wrapped around a json_decode and a prayer. It works until the day the model returns JSON with a stray comma, or sends status as a free-form string when your enum expected one of three values.

Two layers solve this, and they're very different from each other.

Layer 1: format guarantee at decoding time

Structured output with a strict schema isn't "asking the model nicely to return JSON." With strict: true, OpenAI compiles your JSON Schema into a context-free grammar and, at every generated token, masks any token that would violate the schema. The model literally cannot emit output outside the format. In OpenAI's evals, that produced 100% schema adherence.

Notice the difference: JSON mode promises syntactically valid JSON. Strict structured output promises your schema — fields present, types correct, enums respected. It's a parser guarantee, not luck.

response = client.chat.completions.create(
    model="gpt-5",
    messages=[...],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "resposta_suporte",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "resposta": {"type": "string"},
                    "encontrou_no_contexto": {"type": "boolean"},
                    "trechos_usados": {
                        "type": "array",
                        "items": {"type": "string"},
                    },
                },
                "required": [
                    "resposta",
                    "encontrou_no_contexto",
                    "trechos_usados",
                ],
                "additionalProperties": False,
            },
        },
    },
)

Look at the encontrou_no_contexto field. It exists to give you a cheap binary signal: when it comes back false, you don't even render the answer — you trigger the fallback. A guaranteed format gives you somewhere to hang a product decision.

Layer 2: content verification before delivery

The right format with made-up content is still a problem.

The metric that closes that gap is faithfulness: break the answer into atomic claims and check, one by one, whether the retrieved context supports each of them. The score is the supported fraction. It's what frameworks like RAGAS compute, and what teams running evals in production use as a gate — an answer below the threshold doesn't leave the API, it goes down the fallback path.

The architectural point here is the when. This runs before the answer reaches the user, not in a dashboard somebody opens on Monday. A cheap verifier — small model, short prompt, just "is this claim supported by this passage? yes or no" — fits in the latency budget of most products.

And the same dataset works in CI. If you want the continuous eval path before deploy, there's a whole post on building that pipeline.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Source citation: the pattern that cuts customer complaints

Citation does two things at once, and the second is the one that matters to the business.

The first: it improves the answer. When the model has to point to where it got each claim, it goes through the text instead of completing from memory.

The second: it hands the audit to the user. The customer who's suspicious clicks the source and sorts it out on their own. That changes the cost of an error. Without citation, an error becomes a ticket. With citation, an error becomes the user noticing in two seconds that the passage doesn't answer their question.

And you can measure it. Anthropic reports that Endex, by swapping hand-rolled citation for the API's native citations, took source hallucination and formatting issues from 10% to 0%, with 20% more references per answer. In Anthropic's internal evaluations, native citation beats custom implementations by up to 15% in recall accuracy. The mechanism is boring on purpose: the API splits the document into sentences and returns the cited_text, the verbatim passage — not a page number the model found plausible.

Then comes the catch, and it's a big one: a citation existing doesn't prove the citation supports the claim. Recent work on citation grounding evaluated five systems and found 13% to 21% hallucinated citations. It points to the right document, lands on the adjacent paragraph, doesn't support the claim.

So the full pattern has three steps, and the third is the one almost everybody skips:

  1. The model emits the citation.
  2. You programmatically validate that each reference points to a source you actually sent in the context, and strip the ones that don't.
  3. You verify that the cited passage contains the claim. This is where the cheap verifier comes in again.

Step 2 is ten lines of code and catches the fabricated reference. Step 3 is one small-model call and catches the citation that exists but doesn't support the claim. Without both, you have decorative citation.

What will still slip through (and how to fail safely)

None of this gets you to zero. Anyone promising zero is selling something.

The Stanford RegLab study is the best reminder available. Commercial legal research tools built by Thomson Reuters and LexisNexis, with RAG, with citation, with decades of structured proprietary data and sold as "hallucination-free," got it wrong or answered with incorrect grounding on 17% to 33% of queries in a preregistered test with 202 queries evaluated by experts.

Big company. Proprietary data. Serious team. And still.

So the bar for putting the system in front of a paying customer isn't "it doesn't hallucinate." It's four other things.

You know when you don't know. An explicit "not found" signal in the schema, with a defined fallback, is worth more than ten points of average accuracy.

The error has a ceiling. Write down what the system does on its own and what requires a human. Answering a question about business hours: automatic. Stating the amount of a contractual penalty: draft with approval. The yardstick is the cost of the error, not the model's apparent confidence.

The error is traceable. A trace with the retrieved context, the final prompt, the raw output, and the verifier's result. Without that, when the customer complains, you'll be debugging by guesswork.

The error is measured. A dataset of real cases running on every change. Every prompt tweak and every model swap is an untested deploy until you have that.

Failing safely is a product decision made in code. It's not a side effect of a well-written prompt.

Quick FAQ

Does lowering the temperature to 0 fix AI hallucination? No. Temperature 0 makes the output deterministic, not true. If there's no tether to real data, the model completes with the same fabrication, identical every time. You trade random error for reproducible error — which helps with debugging, but fixes nothing.

Do bigger models hallucinate less? Less on average, and not enough for you to stop tethering. The OpenAI paper shows the root is in the training objective and in the way benchmarks score, and neither is solved by scale. A better model raises the floor. The architecture around it is what sets the ceiling.

If my RAG is good, do I still need validation? You do. RAG delivers the right context. Nothing in it guarantees the model used that context. Faithfulness measures exactly the gap between "the right passage was there" and "the answer rests on it."

How much latency does real-time verification cost? Depends on what you verify. Validating references against the sources you sent is pure code, zero cost. Verifying claim support with a small model usually costs hundreds of milliseconds and a fraction of the price of the main call, and you can run it only on the answers the schema flagged as sensitive.

Conclusion

AI hallucination isn't a moral failing of the model. It's the expected behavior of a completer with nothing to hold on to, trained under a regime that rewards guessing.

The fix is architecture, in layers. Grounding pins the answer to your data. Structured output with a strict schema guarantees parseable output and gives you somewhere to hang a decision. Real-time validation knocks down whatever isn't supported before it goes out. Verified citation hands the audit to the user.

None of these layers is optional if the system is going to touch a paying customer. And none of them gets you to zero. What they give you is the thing that actually matters: a system that fails within a limit you chose, tells you when it doesn't know, and leaves a trail so you can fix it.

A pretty prompt doesn't do that. Engineering does.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing