~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tutorials / llm-intent-classification-agent-routing $
Tutorials

LLM Intent Classification: How the Agent Decides Which Route to Take

LS Lucas Souza · · 17 min read
LLM Intent Classification: How the Agent Decides Which Route to Take

"It decides on its own which flow to use." Great. Now show me the log explaining why it picked the wrong flow yesterday at 2 p.m.

Silence.

This is always where the autonomous agent conversation stalls. Routing is the cheapest decision to implement and the most expensive one to debug, because it happens before anything interesting shows up in the trace. The user asked for a refund, the agent walked into the tech support flow, burned six tool calls, and you only found out because the customer complained. The big model didn't make the mistake. A one-word choice made three seconds earlier did.

In this post we crack that piece open: how to do intent classification with an LLM without turning it into a lottery, when the router should be a dumb match and when it should be a model, what to do when no route fits, and how to make every routing decision auditable so you can actually answer that 2 p.m. question.

TL;DR

  • What it is: the layer that classifies the user's intent and decides which specialized flow to send it to, before the agent starts working.
  • Stack/Models: Claude (structured outputs), embeddings for semantic routing, semantic-router or LangGraph for orchestration, OpenTelemetry for the trace.
  • Cost/Access: a deterministic router costs zero, embeddings cost pennies, an LLM router costs one extra call per request. The choice changes the bill at the end of the month.
  • What you walk away with: a router with a closed schema, an explicit fallback policy, trace attributes for auditing, and a way to measure routing accuracy without hand-labeling ten thousand conversations.

Deterministic router vs. LLM router

Anthropic describes routing as one of the basic workflow patterns: "routing classifies an input and directs it to a specialized followup task", and it works well "where there are distinct categories that are better handled separately, and where classification can be handled accurately, either by an LLM or a more traditional classification model/algorithm" (Building effective agents).

Look at the end of that sentence. Either by an LLM, or by a traditional algorithm. The official documentation puts both on the same level, and most teams ignore the second half.

The deterministic router is the one you write with explicit rules: the request came from the /billing endpoint, so route billing. The user clicked the "cancel subscription" button, so route cancelamento. The payload has an order_id, so route pedido. Zero latency, zero cost, one hundred percent auditable, and you can write a unit test for every path.

It doesn't sound sophisticated. That's exactly why it works.

The rule of thumb I use: if the information is already structured at the edge, don't ask the model. Button, HTTP route, form field, attachment type, source channel — that's deterministic signal. Throwing it into a prompt so the LLM can "decide" means paying latency and uncertainty to reconstruct information you already had.

The LLM router comes in when the input is free text and the intent only exists in the semantics. "I bought it last month and it still hasn't arrived" has no order_id, no button, no endpoint. It has intent. That's when you need classification.

And then there's the middle ground almost nobody uses: the embedding router. You embed the query, compare it against examples for each route, and take the closest one. That's what the semantic-router lib does — instead of waiting on slow LLM generation to decide tool use, it decides in vector space and returns a RouteChoice with the route name and similarity score. A survey of production routing patterns puts the embedding router at 16–100ms versus 1–5 seconds for the LLM router.

The math matters at scale. RouteLLM, from LMSYS, trained routers to choose between a strong model and a weak model and reported up to 85% cost reduction on MT Bench while keeping 95% of GPT-4's performance — the best router hit that mark using only 26% of calls to the strong model. Routing isn't architectural decoration. It's the most direct cost lever there is after caching.

Routing is one of the modules we take apart live at AI Engineering Lab 3ª Edição, on September 19 and 20: two days building the entire architecture of a production agent, from tool calling to tracing to cost. If this post made you side-eye your router, that's where we get our hands on it.

The decision criteria, no "it depends"

Input signal Right router
Endpoint, button, form field, channel Deterministic
Free text, few routes (2–8), well separated Embedding
Free text, many routes, or routes that overlap LLM with structured output
Free text + needs to extract a parameter along with it (order_id, date) LLM with structured output

In practice you combine them: a cascade that tries the cheap option first and only escalates to the expensive one when the cheap one isn't confident. Hold on to that idea, it comes back in the fallback section.

LLM intent classification that holds up under ambiguity

Intent classification with an LLM in production (intent classification is the name you'll run into in the libs) is not "ask for the category and run an if on the string". That breaks the first time the model answers "Reembolso" with a capital letter, or "reembolso/troca", or a polite sentence explaining its choice.

Close the schema. With structured outputs the model is forced by constrained decoding to return JSON that matches the schema — it's not "ask nicely and pray", it's a format guarantee.

import json
import anthropic

client = anthropic.Anthropic()

ROTAS = ["reembolso", "status_pedido", "suporte_tecnico", "duvida_produto", "outro"]

SCHEMA = {
    "type": "object",
    "properties": {
        "rota": {"type": "string", "enum": ROTAS},
        "confianca": {"type": "number", "minimum": 0, "maximum": 1},
        "segunda_opcao": {"type": "string", "enum": ROTAS},
        "evidencia": {
            "type": "string",
            "description": "Trecho literal da mensagem que sustenta a rota escolhida.",
        },
    },
    "required": ["rota", "confianca", "segunda_opcao", "evidencia"],
    "additionalProperties": False,
}

def classificar(mensagem: str) -> dict:
    resposta = client.messages.create(
        model="claude-haiku-4-5-20251001",
        max_tokens=300,
        system=(
            "Você classifica a intenção de mensagens de clientes de e-commerce.\n"
            "Escolha UMA rota. Se a mensagem tiver duas intenções, escolha a que "
            "bloqueia o cliente agora e registre a outra em segunda_opcao.\n"
            "Se nenhuma rota descrever a mensagem, use 'outro' com confianca baixa. "
            "Não force uma rota plausível."
        ),
        messages=[{"role": "user", "content": mensagem}],
        output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
    )
    return json.loads(resposta.content[0].text)

Four design decisions in that schema, and none of them is decoration:

Closed enum. The model doesn't invent a new route. If the route list changes, the schema changes with it and the deploy fails loudly instead of failing quietly.

confianca. Without it you have no way to implement fallback or sampling for evals. It's the number that separates "route directly" from "escalate". Careful: LLM self-reported confidence is a signal, not the truth — it ranks well but calibrates poorly. Use it to choose what to review, not to say "92% chance of being right".

segunda_opcao. Real messages carry two intents all the time: "it didn't arrive and I want my money back" is status_pedido and reembolso. Forcing a single choice without recording the alternative throws half of the customer's request in the trash. With the second option recorded, the refund flow can pull up the order status alongside it, and your eval can measure how often the right answer was sitting in second place.

evidencia. A literal excerpt from the message. This field is what turns "the agent decided" into "the agent decided because of this". It's the difference between auditing and guessing.

Notice I used Haiku, not the big model. A router is short-text classification with a tiny output space — it's the cheapest use case there is. Using the expensive model in the router is burning money at the front door; if you want to see the other leaks of the same kind, we've already mapped the most common ones here.

The detail that matters more than the model

Router quality depends less on the model and more on the route definitions. A route with a nice name and no description is guaranteed ambiguity.

Write each route as a three-line contract: what goes in, what explicitly does not go in, and a real boundary example.

suporte_tecnico
  ENTRA: produto recebido que não liga, não conecta, apresenta defeito de funcionamento.
  NÃO ENTRA: produto que não chegou (-> status_pedido).
  NÃO ENTRA: cliente quer devolver produto que funciona (-> reembolso).
  FRONTEIRA: "chegou quebrado" -> suporte_tecnico, porque a ação é troca/garantia.

The NÃO ENTRA line is the one that moves the needle most. You're resolving collisions between routes, and route collision is the root cause of most routing errors — not a lack of model intelligence.

And there's a limit worth knowing about: as the number of options grows, selection accuracy falls off a cliff. The survey cited above shows a drop from ~94% accuracy with 50 tools to 20% with 417, due to attention degradation over long context. If your agent has thirty flat routes, the problem isn't the prompt. It's the topology. Group them into two levels: a domain router with 4–6 outputs, and second-level routers inside each domain.

The fallback: what to do when no route fits

This is where the bug nobody can reproduce in staging lives.

Every router is a total function: it always returns something. If you didn't design the "no route fits" path, the model will fill that void with the closest-looking route — and a lookalike route is worse than no route, because it executes. It opens a ticket, fires an email, hits an API, spends tokens, and delivers a confident answer to the wrong question.

Fallback is a product decision, not a code decision. Three policies, and you have to pick explicitly:

1. Abstain with a disambiguation question. Low confidence, two routes tied: send a short question back to the user. It's the best option in chat and the worst in async processing, where there's nobody around to answer.

2. Generic route with reduced capability. Send it to a flow that only knows how to answer from documentation and escalate. It doesn't execute any action with side effects. This is the safe default.

3. Human escalation. Queue, ticket, human agent. Expensive, which is exactly why it needs a well-calibrated threshold.

The pattern that ties it all together is the cascade: deterministic filter first, embedding router next, LLM as the catch-all for whatever is left that's genuinely ambiguous. The production reference suggests auto-routing above 0.8 confidence, flagging 0.5–0.8 for review, and escalating below 0.5. Treat those numbers as a starting point, not as truth: they depend on your route distribution and on the cost of getting each one wrong.

def rotear(mensagem: str, contexto: dict) -> dict:
    # 1. Determinístico: sinal estruturado ganha sempre.
    if contexto.get("origem") == "botao_cancelar":
        return {"rota": "cancelamento", "metodo": "deterministico", "confianca": 1.0}

    # 2. Embedding: barato, resolve o caso comum.
    rota, score = router_embedding(mensagem)
    if score >= 0.85:
        return {"rota": rota, "metodo": "embedding", "confianca": score}

    # 3. LLM: só o que sobrou.
    decisao = classificar(mensagem)
    decisao["metodo"] = "llm"

    # 4. Fallback explícito. Sem isso, o passo 3 sempre "resolve".
    if decisao["confianca"] < 0.5 or decisao["rota"] == "outro":
        decisao["rota"] = "desambiguacao"
        decisao["metodo"] = "fallback"

    return decisao

One special case of routing we've already covered here is the decision to retrieve or not retrieve: in agentic RAG, the agent decides whether to retrieve context, what to retrieve, and how many times. It's the same mechanism as this section with two routes instead of six — and it's a good place to practice thresholds before applying them to a ten-output router.

It's worth being clear about what fallback is not: it doesn't replace input and output validation. The router picks a path; the thing that blocks malicious content or out-of-policy responses is the guardrails layer. They're different controls, at different points in the pipeline.

▪ Clã Beer and Code

A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.

Join the Clã

Making the decision auditable

Back to the question from the beginning: why did it pick the wrong flow yesterday at 2 p.m.?

You can only answer that if you recorded enough at the moment of the decision. And "enough" isn't the text of the final response — it's the state of the decision. Five fields, on the router span:

  • routing.method — deterministic, embedding, or llm.
  • routing.route — the chosen route.
  • routing.confidence — the score.
  • routing.runner_up — the second option and its score. This is the field that exposes a tie.
  • routing.version — version of the route set and of the classifier prompt.

That last one is the most forgotten and the most valuable. Without a version, when someone complains about an error from three weeks ago you won't know which route definition that decision was made under. And since route definitions change every week, you end up auditing a system that no longer exists.

To avoid inventing your own naming, anchor on the OpenTelemetry GenAI semantic conventions, which standardize attributes under the gen_ai.* namespace and already define spans for agent invocation and tool execution. Worth knowing that these conventions are still pre-stable: in June 2026 the gen_ai.* attributes were moved to a dedicated repository and the names can still change between versions. In other words: anchor on them, but isolate them behind an instrumentation function of your own, so you're not hunting strings across the codebase on the next release.

with tracer.start_as_current_span("route_intent") as span:
    decisao = rotear(mensagem, contexto)
    span.set_attribute("routing.method", decisao["metodo"])
    span.set_attribute("routing.route", decisao["rota"])
    span.set_attribute("routing.confidence", decisao["confianca"])
    span.set_attribute("routing.runner_up", decisao.get("segunda_opcao", ""))
    span.set_attribute("routing.version", ROUTES_VERSION)
    span.set_attribute("routing.evidence", decisao.get("evidencia", "")[:200])

Two caveats from someone who's been burned. First: routing.evidence stores an excerpt of the user's message, so that's personal data flowing into your observability backend — hold it to the same standard as everything else (retention, PII masking, access control). Second: if the router is synchronous in the request path, export the span asynchronously. Auditing can't cost the user latency.

The router is one of the six layers that separate a POC harness from one that survives production — the full anatomy is here, and it's worth reading the other five through the same auditability lens.

Measuring routing accuracy without labeling everything by hand

"Fine, but to know whether the router gets it right I need to label ten thousand conversations."

You don't. You need to label the right ones.

Random labeling is the worst possible use of your time, because most traffic is easy and the router nails it with room to spare. You spend three days confirming the obvious. The strategy that works is stratified sampling by error signal, in four buckets:

Bucket 1 — low confidence. The bottom percentile of confianca. That's where the errors cluster. The uncertainty quantification literature uses exactly this: selecting by the lowest confidence percentile recovers a disproportionate share of the misclassified examples, which makes every label you write worth a lot more.

Bucket 2 — ties. Cases where confianca and the segunda_opcao score are close. A tie is real ambiguity, and real ambiguity usually means your routes collide — the fix tends to be in the definition, not in the model.

Bucket 3 — disagreement between methods. Run the embedding router and the LLM router on the same traffic, in shadow, and label only where they diverge. Disagreement between routers is the cheapest indicator of improvement opportunity there is, and it costs nothing beyond a background call.

Bucket 4 — implicit production signals. This one is free and almost nobody collects it: flow transfers mid-conversation, the user rephrasing the same question, abandonment on the first response, escalation to a human. Every one of those events is a "wrong route" vote with no annotator involved.

With around 200 labeled cases across those four buckets you can already build a useful confusion matrix. And look at it per class, not in aggregate — a router with 92% average accuracy can be sitting at 40% on a low-volume, high-impact route like fraude. The aggregate hides exactly what hurts.

When you want to automate part of the judging, the pattern is LLM-as-a-judge with a rubric and calibration against your human labels — don't trust the judge without measuring the judge. We built that pipeline step by step here, and it plugs straight into the router because the output is categorical and short, which is the easiest case to judge.

Quick FAQ

Should I use the same model as the agent in the router? No. A router is short-text classification with a tiny output — Haiku, a fine-tuned encoder, or embeddings get it done. Save the expensive model for the specialized flow, which is where capability actually shows up. An expensive model in the router is a fixed cost per request with no proportional gain.

Isn't intent routing the same thing as tool calling? No, and confusing the two causes that accuracy degradation with many options. Tool calling picks an action inside a flow; routing picks which flow (with which prompt, which tools, and which permissions) is going to run. Routing reduces the number of tools the agent sees at once, and that's exactly what holds accuracy up.

Do I need a framework for this? No, but it helps you avoid reinventing the graph. In LangGraph the pattern is a schema with Literal via with_structured_output() and add_conditional_edges() wiring the decision to the specialized nodes (docs). You can also do it with a match and three functions. A framework here is convenience, not a requirement.

How do I version route changes without breaking history? Never rename or recycle a route's identifier. The old route gets deprecated, the new route gets a new id, and routing.version goes up. That way the trace from three weeks ago stays interpretable and your eval doesn't mix two different definitions into the same number.

Conclusion

Routing looks like the trivial part of the agent and it's the one that pays back the most per line of code: it cuts cost by choosing the right model, cuts errors by shrinking each flow's decision space, and cuts debugging time when you record the why behind the choice.

Three things to take with you: structured signal doesn't go to the LLM, fallback is a product decision and has to exist before the bug shows up, and a decision without routing.version in the trace is a decision you can't audit later.

The natural next step is to look at the router inside the full design — it's a layer, not a system. The six-layer agent architecture blueprint shows where it fits among context, tools, guardrails, and observability.

So, can you answer what happened yesterday at 2 p.m.?

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing