Half of the "AI projects" I reviewed in 2026 could have been solved with an if and a SELECT. The other half became agents because the team wanted to say they had AI.
That's not a LinkedIn joke. It's the pattern. A request comes in, someone says "this is an agent use case," and six weeks later there's an orchestrator with four tools, a vector store, a reflection loop and a bug nobody can reproduce. The original problem was sorting tickets into six queues. Nobody stopped for three minutes to ask when to use an AI agent for real and when anything else would have done the job.
This post is the handbrake. Four questions, three minutes, and at the end you know whether what's in front of you calls for an if, a query, a prompt, a deterministic flow or an actual agent. With five real cases and the verdict on each one.
TL;DR
- What it is: a 4-question test to pick the smallest mechanism that solves the problem, before you write the first line.
- The ladder:
if→ query → prompt → deterministic flow → agent. You only step down a rung when the one above fails. - The house rule (and Anthropic's): "we recommend finding the simplest solution possible, and only increasing complexity when needed" (Building effective agents).
- The scoreboard: out of 5 real cases that showed up labeled "agent," 1 was an agent.
- The cost of overshooting: Gartner projects that over 40% of agentic AI projects will be canceled by the end of 2027 due to cost, unclear business value or inadequate risk controls.
When to use an AI agent: the 4 questions of the 3-minute test
Four questions. About 45 seconds each. If it took you more than three minutes, you're not testing the problem: you're building a case to justify what you already wanted to do.
You answer in order. The first yes ends the test and picks the mechanism. Don't keep going down out of curiosity.
Q1 — Does the rule fit in an if?
The operational question: is the decision criterion enumerable, stable and written down somewhere?
If the conditions live in a contract, an internal policy, a table of brackets or a tax rule, they're deterministic by definition. Collections after 5 days overdue. Free shipping over R$ 199. Account lock at 3 chargebacks. That's not a "complex decision": it's a business rule, and an auditable business rule has to live in code you can read in a PR.
If the answer is yes, you're done. A well-named if with a unit test is more reliable, cheaper and faster to change than any prompt.
Q2 — Is the answer already in your database?
If what the user wants is retrieval or aggregation of structured data you already have, the mechanism is a query.
"How much did we sell per region last quarter" is not an AI problem. It's a GROUP BY. What usually happens is the team confuses the question in natural language with the answer: the hard part is producing the number, and the number is in Postgres. The natural language, if it's even necessary, is a thin shell on top — and a thin shell doesn't justify an agent.
If the answer is yes, you're done. Query, cache and an endpoint.
Q3 — Is it a language transformation, one shot?
Classify. Extract. Summarize. Rewrite. Translate. Compare two texts.
If the task is to take an ambiguous piece of text, produce an output and that's it, you need one LLM call with a closed schema. Not an agent. Anthropic itself is explicit: "for many applications, optimizing single LLM calls with retrieval and in-context examples is usually enough."
One call. Input, output, done. If the result comes out crooked, the problem is almost always context engineering, not a lack of autonomy. And if what scares you is the model returning broken JSON, that fear has a known fix: structured output with a schema, not a defensive parser.
If the answer is yes, you're done. A versioned prompt, with an eval, and you sleep at night.
Q4 — Does the path change at runtime?
This is the question that separates a flow from an agent, and it's the only one that really matters.
Grab a sheet of paper. Try to draw the entire execution graph before running it: step 1, step 2, this branch if the score is low, that other one if the customer is enterprise. Managed to draw all of it? Then it's a deterministic flow. An LLM can sit in three boxes of that graph and it's still a flow — because the thing deciding the order is your code.
It's only an agent when the next step depends on what the model found out in the previous step, and you can't enumerate the trajectories ahead of time. That's Anthropic's operational definition: "workflows are systems where LLMs and tools are orchestrated through predefined code paths; agents are systems where LLMs dynamically direct their own processes and tool usage."
And there's a second filter glued to this one, which almost everybody skips: can you automatically verify that it worked? An agent with no verification signal isn't autonomy, it's a bet with the company's money. If there's no test, invariant check, API confirmation or human in the loop able to say "this result is correct," you don't have an agent — you have a generator of work for somebody else.
The test is just the entry filter; once the problem genuinely passes it, the hard part begins, which is tool calling, structured output, routing, memory, grounding, tracing, evals and making the cost math work. That's the list we open up live at the AI Engineering Lab 3rd Edition, September 19 and 20, from 9 a.m. to 1 p.m., online: two days of agent architecture with the error showing up on screen and getting debugged on the spot.
if vs. query vs. prompt vs. flow vs. agent
The ladder, with the cutoff criterion for each rung:
| Mechanism | When it's the right answer | Cost per run | What you need to operate it |
|---|---|---|---|
if / rule |
Enumerable, stable criterion, written in a policy or contract | ~zero | Unit test |
| Query | Structured data you already have | ~zero | Index and cache |
| 1 prompt | Language transformation, input → output, no state | 1 call | Closed schema + sample eval |
| Deterministic flow | Several steps, but the graph fits on paper before you run it | N calls, N known | Orchestration, retry, idempotency |
| Agent | Next step depends on what the model found out, and the result can be verified | N calls, N variable | Tracing, evals, guardrails, budget, fallback |
Look at the last column. That's the real price, and it doesn't show up in anyone's estimate. It's what settles the deterministic automation vs. AI agent fight in practice, not the elegance of the architecture.
An if has two paths. You test it in five minutes and never think about it again.
An agent has a trajectory space. You don't test "the result": you test the behavior under variation, which takes a dataset of cases, a rubric, a judge, a trace per run and a spending cap. That's not a picky engineer being precious — it's what separates an agent that runs in production from a demo that wowed the Tuesday meeting.
And there's the middle ground almost nobody considers: the deterministic flow with an LLM inside. You fix the order of the steps in code and use the model only where there's language ambiguity. Then, if at some point the flow has to pick between three paths, that becomes a problem of intent classification with an explicit fallback policy — which is a whole rung cheaper and more auditable than handing the decision to the loop.
5 real cases and the verdict on each one
All of them showed up labeled "agent." Names changed, scopes preserved.
Case 1 — "Ticket triage agent"
The ask: classify a support ticket into one of six queues and assign a priority.
The test: Q1 no (the criterion is free-text content, not a rule). Q2 no. Q3 yes — it's a language transformation, one shot, with a closed output.
Verdict: one LLM call with a schema. No loop, no tool, no memory.
// 1 chamada. Enum fechado com 6 filas + "indefinido" como saída de escape.
$saida = $llm->structured(
schema: TriagemSchema::class, // fila: enum, prioridade: enum, confianca: float
system: $this->politicaDeFilas, // a política, versionada, não improvisada no prompt
input : $ticket->corpo,
);
if ($saida->confianca < 0.7 || $saida->fila === 'indefinido') {
return $this->paraFilaHumana($ticket); // fallback explícito, não silencioso
}
$ticket->update(['fila' => $saida->fila, 'prioridade' => $saida->prioridade]);
Look at the if on the fallback line. That's what makes the system operable: you have a number (confianca) and a declared escape route. That same number becomes the metric that decides, months later, whether this ever needs to turn into something else.
Case 2 — "Monthly report agent"
The ask: a monthly sales summary in the leadership team's inbox, "with AI analyzing the data."
The test: Q1 no. Q2 yes — every number requested was in three database tables.
Verdict: query + template. The only piece left for AI was the prose read-out paragraph at the top of the email, and even that one gets the numbers already computed by SQL, ready to go, in the prompt. The model writes; it doesn't count.
That separation isn't nitpicking. A language model botching arithmetic in a leadership report is the fastest way to torch AI's credibility inside the company for two years. Leave the sum to Postgres.
Case 3 — "Collections agent"
The ask: decide who to collect from, when to collect and through which channel.
The test: Q1 yes, on the first question. The rules were written in the customer's contract: overdue brackets, allowed channels, time windows, exceptions by plan type.
Verdict: if + cron. And here the argument isn't even about cost, it's about risk. Collections carries legal consequences and has to be reproducible: someone, someday, is going to ask why that customer got that message on that day. "The model decided" is not an answer that survives an audit.
A rule with legal effect lives in deterministic code. Always.
Case 4 — "Onboarding agent"
The ask: answer a new hire's questions about the internal handbook, the HR policies and the reimbursement process.
The test: Q1 no. Q2 no (the answer is in a PDF, not a table). Q3 almost, but the content doesn't fit in the prompt. Q4 no — the graph fits on paper: retrieve passages → assemble context → answer citing the source → if nothing was found, say nothing was found.
Verdict: a three-step deterministic flow, with RAG. Three fixed steps. Zero autonomy. And the real win here isn't "intelligence": it's grounding, which is what actually cuts hallucination by forcing the answer to cite the source document.
By the way, before you assume you need RAG, it's worth running through the criteria for choosing between fine-tuning, RAG and prompt — the wrong rung on that choice costs weeks too.
Case 5 — "Refund agent" (this one passed)
The ask: resolve refund requests end to end. Read the request, look up the applicable policy, check the order history, verify whether the product came back, decide, execute the refund and notify the customer.
The test: Q1 no (there is a policy, but it has exceptions that depend on context: carrier delay, repeat customer, item condition). Q2 no. Q3 no (it's several steps with state). Q4 yes — the order of the checks depends on what turns up: if tracking shows the package was lost, it's one path; if the product was delivered and refused, it's another; if the customer already had two refunds this quarter, it's another. And, crucially, it can be verified: the refund goes through or it doesn't, the amount matches or it doesn't, the cited policy exists or it doesn't.
Verdict: it really is an agent. One out of five.
And this is where the conversation grows up, because an agent isn't the end of the work, it's the beginning. If you've reached this rung, it's worth revisiting the honest definition of what an AI agent is — model, tools, memory and orchestration, the four blocks — and then deciding how much of the stack (LangGraph, Mem0, LangChain, MCP) you actually need, which tends to be a lot less than Twitter suggests.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãThe cost of overshooting (and why nobody measures it)
The question of when not to use AI is expensive precisely because it's asymmetric.
Undershooting is visible. You ship a dumb rule, the user complains in the first week, you fix it. It hurts fast and cheap.
Overshooting is invisible. The agent works. It demos well. It impresses at the sprint review. And the cost shows up slowly, in four layers nobody put in the estimate:
Latency. An if answers in microseconds. An agent chains sequential calls, and the p95 isn't the sum of the averages — it's the sum of the tails. Anthropic says it bluntly: "agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense."
Cost per run. The loop resends the history at every step. What was one call becomes eight, and each one carries the accumulated context of the ones before it. That's why an agent's token bill doesn't grow linearly with usage — and why the real invoice deserves to be opened up layer by layer before it turns into an end-of-month surprise.
Operational surface. Tracing per run. A dataset of cases. An eval with a rubric. A spending guardrail. An idempotent retry policy. Human fallback. None of that is optional in production, and none of it was on the Jira card that said "build agent for X."
Cost of change. This is the worst one. Changing an if is a three-line PR with a test. Changing an agent's behavior means touching the prompt, re-evaluating the entire dataset, finding out the change improved two cases and broke five others, and negotiating which regression is acceptable. You traded "edit a rule" for "renegotiate a statistical behavior."
Nobody measures it because there's no alert for it. There's no "unnecessary complexity" metric in Datadog. The aggregate symptom shows up far away: Gartner estimates that only about 130 of the thousands of vendors calling themselves agentic are real, and calls the rest agent washing — rebranding of assistants, RPA and chatbots with no actual agentic capability. On the buyer's side, the MIT NANDA report on the state of AI in business in 2025 found about 95% of organizations with no measurable business return, despite tens of billions invested.
Don't read that as "AI doesn't work." Read it as what it is: most of the money was spent one or two rungs further down the ladder than necessary.
What to do when the pressure comes from above
None of this solves the political problem. You run the test, it comes out if, and the director wants to be able to say the company has AI. That's real, and there's no point pretending it isn't.
Three moves that work better than a technical lecture.
1. Don't fight the word. Fight the scope.
Arguing over whether "this is an agent or not" is a naming fight you won't win and don't need to win. Deliver the result that was asked for, in the smallest mechanism that solves it, and let the business side call it whatever they want. If saying "intelligent triage assistant" buys three weeks of peace and avoids six of overengineering, that's a good deal.
2. Ship the deterministic version first, instrumented.
One sprint, not six weeks. Deterministic, with an explicit fallback and a log of everything it couldn't resolve. That gives you two things at once: value in production now, and the data that will decide the next step — instead of opinion.
-- A métrica que decide se um dia isso vira agente.
-- Fallback alto e subindo = o grafo fixo não cobre mais a realidade.
SELECT
date_trunc('week', criado_em) AS semana,
count(*) AS total,
count(*) FILTER (WHERE resultado = 'fallback_humano') AS caiu_pra_humano,
round(100.0 * count(*) FILTER (WHERE resultado = 'fallback_humano')
/ nullif(count(*), 0), 1) AS pct_fallback
FROM execucoes_triagem
WHERE criado_em > now() - interval '90 days'
GROUP BY 1
ORDER BY 1 DESC;
3. Write the cutoff line before, not after.
Record it in an ADR, on the wiki, in the README, wherever — but in writing, with a number and a date. Something like this:
## ADR-014: triagem de tickets sem agente
Decisão: 1 chamada de LLM com schema fechado + fallback humano.
Alternativa descartada: agente com tools de CRM e busca.
Vira agente quando (qualquer um):
- pct_fallback > 20% por 3 semanas seguidas; OU
- o fluxo passar a exigir ação em sistema externo cujo caminho
não dá pra enumerar (ex.: abrir exceção de política caso a caso).
Revisão: 01/12/2026.
That changes the nature of the conversation. You stop being "the dev who doesn't want to do AI" and become the one who defined the objective criterion for when to do it. It's the same logic as choosing a way of working by context instead of by hype, which we already broke down in the decision table for agentic code, vibe coding and SDD.
Oh, and a common-sense warning: none of this counts as an excuse not to learn. The reason not to build an agent here is that this problem doesn't call for one. When the problem that does call for one shows up — and it does — you need to know how to build it. Those are two different skills: knowing how to decide and knowing how to execute. Missing either one is expensive.
Quick FAQ
If I already have the agent built, do I throw it away?
Not necessarily. Measure first: p95 latency, cost per resolved run and accuracy against a dataset of real cases. Then compare with the equivalent deterministic version. If the agent doesn't win on any of the three, the refactor to a fixed flow usually pays for itself in a quarter — and you reuse the prompts and the tools.
Where does the "deterministic flow with an LLM" fit in? Isn't that half an agent?
No. The difference is who decides the order. In the flow, it's your code: the graph is written down and you can draw it before running it. In the agent, it's the model, at runtime. A flow can have five LLM calls and still be deterministic from a control standpoint.
My problem has free-text input. Doesn't that already force me to use an agent?
No. Free text forces you to use an LLM, which is the Q3 rung. Ambiguous input and a variable execution path are independent things: you can have natural language on the way in and a completely fixed graph after it. That's exactly the ticket triage case.
How do I explain this test to the team without coming off as the "no" guy?
Flip the framing. It's not "let's avoid AI": it's "let's spend the complexity budget where it pays off." Every team has a ceiling on the complexity it can operate well. Burning that ceiling on a problem that was an if means having no gas left for the problem that really was an agent.
Conclusion
Four questions. if, query, prompt, flow, agent. The first yes ends it.
Out of five cases that showed up with the word "agent" in the card title, one was an agent. That ratio isn't just my anecdote — it's the same direction Gartner and MIT point to with a large sample: money spent one or two rungs further down the ladder than necessary.
The sign of maturity in AI engineering isn't being able to build an agent. Anyone can put together a loop with tools over a weekend. The sign is looking at a problem and being able to say, in three minutes and without fooling yourself, that it doesn't need one.
And when it does — because it will — you already know where to look for the next step: what an AI agent actually is and, when it's time to write the code, which way of working fits your context.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã