~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / llm-structured-output-json-that-doesnt-break $
Tutorials

JSON That Doesn't Break: LLM Structured Output Without a Defensive Parser

LS Lucas Souza · · 16 min read
JSON That Doesn't Break: LLM Structured Output Without a Defensive Parser

If your code has a try/catch wrapped around a json_decode of the model's response, you don't have a contract. You have hope.

And hope doesn't scale. It works in your ten manual tests, gets through code review because "that's just how LLMs are," and breaks on a Tuesday at 2 p.m. when a customer sends an input that pushes the response up against the token limit. Then the JSON arrives half-finished, the parser blows up, and the job lands in the failed queue with nobody understanding why.

The good news: structured output in LLMs stopped being a prompt problem and became an architecture problem, with an engineering solution. In this post you'll see why the model breaks JSON, what changes when you guarantee the AI's structured output in the decoder itself instead of asking nicely, how to design a schema the model can actually satisfy, and how to build validation and retry that don't turn into an infinite loop.

TL;DR

  • What it is: guaranteeing that the LLM's output arrives as a typed, validatable structure, not as text that sometimes looks like JSON.
  • How it's done: constrained decoding at the provider (output_config on Anthropic, text.format with strict: true on OpenAI, responseSchema on Gemini) + semantic validation in your own layer + retry with a fixed budget.
  • Cost/Access: available on the paid APIs of the big three. It charges extra system prompt tokens and adds latency only on the first request for each schema.
  • Scope: this post is about the model's output. Designing the input schema for tools is covered in Tool use in practice.

Why the model breaks JSON (and when)

First, drop the idea that the model "gets JSON wrong" at random. It fails in patterns, and patterns can be handled.

The most common mode is the model being too helpful: it returns the right object, but wrapped in a markdown fence, with a "Sure! Here's the JSON you asked for:" in front. Then comes the trailing comma on the last array item. Then the unescaped quote inside a free-text field. Then the field you never defined, invented because it seemed useful. Then the enum with a value outside the domain, like "status": "pendente_aprovacao" when your contract only accepts pending, approved and rejected.

And then there's the failure that hurts most in production, because it isn't the model's fault: truncation. If generation hits the token ceiling before closing the object, you get half a JSON. On OpenAI this comes back with status: "incomplete" and incomplete_details.reason: "max_output_tokens" — an explicit signal your code probably ignores, because it went straight to parsing the body. There's also the safety refusal case, where the API returns a refusal field instead of following the schema.

The root cause is simple and worth internalizing: without a constraint in the decoder, format is just one more statistical preference. You wrote "respond in JSON" in the prompt and the model learned that, after that instruction, JSON tokens are likely. Likely. Not mandatory. Sampling is probabilistic, and long tails exist. The bigger the schema, the longer the free-text field and the more unusual the input, the fatter that tail gets.

This is exactly the kind of decision that separates a prototype from a production system, and it's one of the modules we break down live at the AI Engineering Lab 3rd Edition, on September 19 and 20, alongside tool calling, routing, memory, grounding and tracing. Two days of online immersion, from 9 a.m. to 1 p.m., with running code instead of slides.

Structured output in LLMs vs. "respond in JSON, please"

There are three levels of guarantee, and most teams are on the first one thinking they're on the third.

Level 1 — asking in the prompt. Zero guarantee. You're negotiating with a probabilistic system. It works well enough to fool you in development and badly enough to wake you up in the middle of the night.

Level 2 — JSON mode. OpenAI's response_format: { type: "json_object" } guarantees the output is syntactically valid JSON. That's it. Nothing stops the model from returning {"foo": "bar"} when you expected fifteen fields. Syntax guaranteed, contract not.

Level 3 — structured outputs. Here the provider compiles your JSON Schema into a grammar and constrains sampling token by token. Tokens that would violate the schema simply aren't candidates. It's not the model "trying to comply": it's the decoder preventing invalid output from existing.

On Anthropic, the parameter is output_config:

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": email_bruto}],
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "nome": {"type": "string"},
                    "email": {"type": "string", "format": "email"},
                    "plano": {"type": "string", "enum": ["free", "pro", "enterprise"]},
                    "quer_demo": {"type": "boolean"},
                },
                "required": ["nome", "email", "plano", "quer_demo"],
                "additionalProperties": False,
            },
        }
    },
)

Watch out for a migration gotcha: the old output_format parameter is still accepted by the API during the transition, but the Python SDK v1.0+ raises TypeError if you pass it to client.beta.messages.create(). Use output_config. The feature entered public beta in November 2025 and went GA in February 2026.

On OpenAI, the format is text.format with strict: true:

text={
    "format": {
        "type": "json_schema",
        "name": "extracao_lead",
        "strict": True,
        "schema": schema,
    }
}

On Gemini, you combine responseMimeType: "application/json" with responseSchema, and there's an extra worth knowing about: propertyOrdering, which pins the order in which the model emits the fields. Standard JSON Schema ignores key order; autoregressive generation doesn't. That matters more than it seems, and I'll come back to it.

One split that confuses a lot of people: a tool schema is input, a response schema is output. When you declare a tool with strict: true, you're guaranteeing the format of the arguments the model sends to your function. That's the subject of the post on designing tools the LLM can actually use. Here we're on the other side of the flow: what the model returns to your system at the end of the conversation. Both use JSON Schema, both have a strict mode, and they combine: it's the same function calling schema you already write, pointed at opposite ends of the flow. But they're different contracts, at different points in the architecture.

And structured output isn't free. It injects an additional system prompt explaining the format, which raises the input token count. The first request with a new schema pays for grammar compilation (somewhere between 100 and 300ms on Anthropic), and the result is cached for 24 hours from last use. Changing the schema invalidates that cache. So does changing the set of tools. And touching output_config.format invalidates your prompt cache. Hold on to that: it turns into a deploy decision later on.

Designing a schema the model can satisfy

Here's the expensive mistake: assuming structured output accepts all of JSON Schema. It doesn't. Each provider implements a subset, and whatever falls outside isn't silently ignored — it usually becomes a 400 error at deploy.

On Anthropic, what's left out is recursive schemas, external $ref, numeric constraints (minimum, maximum, multipleOf), string constraints (minLength, maxLength) and additionalProperties with any value other than false. What gets through is enum, const, anyOf, internal $ref, the usual string formats (date-time, email, uri, uuid) and minItems only with 0 or 1.

On OpenAI, strict mode is even more rigid on one specific point: every field has to be in required. There's no such thing as an optional field. If you want optionality, you emulate it with a type union:

{
  "cnpj": { "type": ["string", "null"] }
}

And there are ceilings: around 100 properties across the whole schema and nesting depth up to 5. The root object can't be anyOf, and default isn't supported.

On Gemini, schemas that are too large or too nested get rejected with a 400 and a message that rarely points to the offending field.

With that in hand, four practical design rules:

1. Flatten. Deep nesting is where everything breaks: it hits the provider's ceiling, eats tokens, and raises the odds of the model losing the thread in the middle of an object. If you have a cliente.endereco.geo.coordenadas.lat, you can probably have cliente_lat.

2. Close the domain with enum. Every field that in practice has five possible values should be an enum, not a free string. It's the difference between validating afterward and being unable to generate it wrong. Enum is the most underused tool in LLM schemas.

3. Use description as a micro-prompt. Each property's description field goes to the model along with everything else. {"prazo_dias": {"type": "integer", "description": "Dias úteis até a entrega. Se o texto não disser, use 0."}} resolves ambiguity at the point where it happens, and it's far more effective than another paragraph in the system prompt.

4. Order the fields to think before deciding. Generation is autoregressive: what comes out first conditions what comes next. If the object starts with "decisao": "aprovado" and only then has "justificativa", the model decided before reasoning and the justification became rationalization. Flip it. Put a short reasoning field before the decision fields.

That fourth rule is backed by the literature, and it's worth knowing the full counterargument. The paper Let Me Speak Freely? showed that format restriction degrades the model's reasoning ability, and that the stricter the restriction, the bigger the drop. On classification and field-filling tasks the rigid format helps. On tasks that require real reasoning, it gets in the way. The recommended mitigation is the same thing rule 4 does in practice: let the reasoning happen loose first, and only then tighten the structure. Whether that's two fields in the same object or two separate calls.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

Validation and retry without an infinite loop

Now the part almost everyone skips: a valid schema is not correct data.

The decoder guarantees that {"cnpj": "00000000000000"} matches the contract. It doesn't know that CNPJ doesn't exist. It guarantees that {"prazo_dias": -5} is an integer — and since numeric constraints aren't supported in strict mode on Anthropic, the negative sails right through. Constrained decoding eliminates the class of errors of form. The class of errors of content is still fully intact, and it's yours.

So it's two layers, always:

from typing import Literal
from pydantic import BaseModel, Field, field_validator

class Lead(BaseModel):
    raciocinio: str = Field(description="Por que classificou assim. 1 frase.")
    nome: str
    email: str
    plano: Literal["free", "pro", "enterprise"]
    prazo_dias: int

    @field_validator("prazo_dias")
    @classmethod
    def prazo_positivo(cls, v: int) -> int:
        if v < 0:
            raise ValueError("prazo_dias não pode ser negativo")
        return v

And the retry. The pattern that works has four properties, and the order matters:

MAX_TENTATIVAS = 2

def extrair(texto: str) -> Lead:
    erro_anterior = None
    for tentativa in range(MAX_TENTATIVAS + 1):
        mensagens = montar_mensagens(texto, erro_anterior)
        bruto = chamar_modelo(mensagens)          # já com structured output
        try:
            return Lead.model_validate_json(bruto)
        except ValidationError as e:
            erro_anterior = e.json()               # devolve o erro pro modelo
            metrics.increment("lead.retry", tags=[f"tentativa:{tentativa}"])
    raise ContratoQuebrado(texto, erro_anterior)   # dead letter, não loop

First: fixed budget. Two extra attempts, max. Infinite retry with an LLM isn't resilience, it's burning credits with pretty logs. If it failed three times, the problem isn't luck.

Second: send the validation error back to the model. The Pydantic message says exactly which field failed and why. That's extremely high-value context and it costs almost nothing. Blind retry, repeating the same prompt, tends to produce the same failure.

Third: a terminal failure is a dead letter, not a swallowed exception. The record goes to a review queue with the original input and the last error. Someone looks at it. Nobody pretends it didn't happen.

Fourth: retry rate is a product metric, not log noise. If one specific field drives 8% of retries, the schema is badly designed — ambiguous description, incomplete enum, a field that should be optional. The metric tells you where to fix the contract. It's the same reasoning behind building honest evals for your agent: without measuring, you're in the dark.

And a boundary rule: the validated object is the only thing that crosses into the rest of the system. No raw LLM array floating through three service layers. In Laravel that's a DTO at the edge, and from there inward it's normal code:

public function handle(array $payloadDoModelo): LeadData
{
    $validado = Validator::make($payloadDoModelo, [
        'nome'       => ['required', 'string', 'max:120'],
        'email'      => ['required', 'email'],
        'plano'      => ['required', Rule::in(['free', 'pro', 'enterprise'])],
        'prazo_dias' => ['required', 'integer', 'min:0'],
    ])->validate();

    return LeadData::from($validado);
}

Redundant with the schema? Partly. And that's on purpose. The schema protects against the model; the validator protects against the schema being wrong, against a provider swap, against the day someone touches the prompt and forgets about the contract.

Versioning the contract when the schema changes

An output schema is an internal API. And an internal API that changes without a version breaks its consumers.

The moment this bites is predictable: you add a field, deploy, and the records written yesterday don't have that field. Or you rename status to situacao and the dashboard that reads the JSON column loses half its data. None of that is an AI problem. It's a botched migration, with an LLM in the middle.

Three practices that fix it:

Store the version alongside the data. A schema_version field in the table, filled in by your code (not by the model). When you need to reprocess or debug a record from three months ago, you know which contract it was generated against.

Classify the change before making it. Adding a field with a default value on your side is additive and cheap. Removing a field, renaming a field, tightening an enum or changing a type is breaking — and requires running v1 and v2 side by side until the consumer migrates, exactly as you would with a REST endpoint.

Lock it down with a golden set in CI. Twenty to fifty real inputs, including the ugly ones: empty field, text in two languages, the PDF that came in with bad OCR, the case that caused last month's incident. On every schema or prompt change, run them all and compare. It's the same principle as adversarial cases to find out where the prompt breaks, applied to the contract instead of the text.

And remember the operational cost I mentioned earlier: changing the schema invalidates the provider's grammar cache and your prompt cache. In the first window after the deploy you'll see higher latency and token cost. It's not a regression. It's the bill for changing the contract — but it's good to know that before the alert fires, not after.

Limitations and things to watch

Structure charges a reasoning tax. That's the conclusion of "Let Me Speak Freely?" and it bears repeating: don't force a tight schema on top of a task that requires thinking. Separate the steps.

Coverage is uneven. Smaller models, open-source models you serve yourself and some compatibility endpoints don't implement real constrained decoding. If your architecture has fallback between providers, test the fallback with the same golden set. Finding out mid-incident that the backup model only has JSON mode is bad.

Truncation is still your enemy. Structured output doesn't guarantee the response fits in max_tokens. Check the stop reason before parsing, always. A large object with free-text fields is a natural candidate to overflow.

The schema is a data surface. Field names and descriptions go to the provider on every request. Don't put sensitive business rules, customer names or confidential internal structure inside description without thinking.

Quick FAQ

Does structured output replace validation in my backend? No. It guarantees form, not content. A nonexistent CNPJ, a date in the past, a value out of range and the ID of a record that doesn't exist all pass cleanly through the schema. Keep both layers.

Why does my valid schema return a 400? It's almost always a keyword outside the supported subset. The usual suspects: minLength/maxLength, minimum/maximum, a recursive schema, additionalProperties other than false, or a field missing from required in OpenAI's strict mode. Compare your schema against the provider's support table before blaming the model.

Does structured output make the response slower or more expensive? A bit of both. It injects an additional system prompt and the first request with each new schema pays for grammar compilation. After that the grammar is cached for 24 hours and the overhead disappears. In practice, it's cheaper than the retry you're doing today.

Do I use strict: true on the tool or a schema on the response? Both, at different points. Strict mode on the tool guarantees the arguments that reach your function; a schema on the response guarantees the final object that goes into your database. One handles input, the other output.

The contract is yours, not the model's

The summary fits in one sentence: stop treating LLM output as text and start treating it as an interface. The schema constrains the decoder, semantic validation handles what the schema can't reach, retry has a budget and sends the error back to the model, and the contract is versioned like any other API.

What changes in your code is modest and the effect is big: the defensive parser goes away, the generic try/catch goes away, and the model's response starts entering your internal systems like any other typed payload. The LLM becomes one more service with a contract, not an exception everyone handles out of fear.

And if you haven't designed the tools on the other side of the flow yet, that's the natural next step: tool use in practice, designing tools the LLM can actually use.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing