Lucas Souza
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
Claude went down today and Codex went with it. On the other side, Anthropic engineers with the same frozen terminal. Do they have private servers? Do they switch to another model? Do they go back to reading logs by hand? The answer, cross-referencing a reliability team talk, leaked code, postmortems and 358 status page incidents.
Tool lists rot in twelve months, yours included. That's why each of the ten comes with a swap criterion: the ten slots in the AI engineer's stack in 2026, the default for each one, and the objective signal that tells you when to rip it out. With data from the 2026 OWASP Top 10 and the Pragmatic Engineer survey.
Half of the "AI projects" I reviewed in 2026 could have been solved with an if and a SELECT. The other half became agents because the team wanted to say they had AI. Here's the 3-minute test: 4 questions that decide whether your problem calls for an if, a query, a prompt, a deterministic flow or an actual agent. With 5 real cases, the verdict on each one (spoiler: 1 in 5 was an agent) and the invisible cost of overshooting.
The agent decides on its own which flow to use, but nobody can explain why it picked the wrong one yesterday at 2 p.m. Deterministic router vs. LLM router, intent classification with a closed schema and evidence, an explicit fallback policy, trace attributes to audit the decision, and how to measure routing accuracy without hand-labeling ten thousand conversations.
If your code has a try/catch wrapped around a json_decode of the model's response, you don't have a contract. You have hope. How to get truly guaranteed structured output from an LLM: constrained decoding at the provider, a schema the model can satisfy, semantic validation and retry with a fixed budget, with no infinite loop and no defensive parser.
Anthropic released Claude Fable 5.1 on September 1, 2026. It beats Opus 5 on every published benchmark, but almost always by 2 to 3 points, and it costs twice as much per token. The official numbers, the real math on an agentic session with a warm cache (where the cost ratio drops from 2x to 1.3x), the decision tree between Fable 5.1, Opus 5 and Sonnet 5, and the 3 breaking changes that break your code if you just swap the model id.
Your prototype cost R$ 12 on day one. In production, with 200 users, it costs almost R$ 4,000 a month. We open up the token bill of a real agent layer by layer (system prompt, tool definitions, RAG chunks, history, tool results) and show where almost 80% of the money disappears. In the end, prompt compression, history pruning, and context selection cut the bill by 73%, with the before-and-after spreadsheet.
The threads say Claude Opus 5 regressed, and Google already answers yes. But the most likely explanation isn't a model nerf: Anthropic cut more than 80% of Claude Code's built-in system prompt for the Claude 5 generation. The restraint defaults are gone, and the responsibility moved to your CLAUDE.md. What you can measure, what's perception, and how to tame it without switching vendors.
A class action in California (Kahn v. Anthropic, 3:26-cv-05763) alleges that the Claude Max 20x plan delivers 6x to 8x the usage of Pro, not 20x. What the complaint backs up with internal documents, why the multiplier only applies to the 5-hour session while the weekly cap is what actually locks you out, and how to find out which of the two you're hitting before you renew.
Grounding isn't RAG, it's the property you want: the answer pinned to data that's yours. The techniques that tether the LLM, structured output with a strict schema, the faithfulness check that runs before delivery to the paying customer, and what will still slip through after all of that.
An honest map of the agent orchestration and memory ecosystem. What LangGraph, Mem0, LangChain and MCP actually do, the point where each one stops being overhead, and the 5-question decision tree to run before installing anything.
The third LLM infrastructure provider already has a São Paulo region, a catalog with Grok 4.3, Gemini 2.5, Llama 4, and Cohere Command A, and bills per character instead of per token. The real Oracle Generative AI catalog, how to call the service with the OpenAI SDK, the dedicated cluster math, and the objective criteria for deciding between OCI and Bedrock.