Lucas Souza
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
The decision most people get wrong by overshooting. Objective criteria for choosing between prompt, RAG, and fine tuning, what each path costs with real numbers, the three rare cases where training a model wins, and the five-minute test that separates a knowledge problem from a retrieval problem. Plus the news that changes the math: OpenAI is shutting down its fine-tuning platform.
The operational fundamentals of LLMs with no parrot analogy: what the model does at each token, why it's stateless, how the context window degrades long before the limit, and why temperature 0 isn't deterministic. Each concept closes with the architectural consequence it forces you to build, with the official Anthropic and OpenAI docs, the Lost in the Middle paper, and Chroma's context rot study in hand.
Half the prompt techniques disappeared because the model learned them on its own. The other half became API parameters. What survived, what turned into folklore, and why in a real system the problem stopped being the sentence and became what goes into the context window.
A twelve-word post on X became a paradigm with its own paper in five weeks. I went and read the survey looking for the benchmark that justifies the new rung: there isn't one. What separates hype from engineering in graph engineering, with the numbers traced back to the source, which ones come from vendors and which are independent, and the five-question yardstick for deciding whether your case calls for a graph or you just want the new badge.
AWS documents how to turn on Bedrock really well. Nobody documents the rest: the difference between CloudTrail and model invocation logging, the fact that São Paulo doesn't give you data residency, and what shows up on the bill in reais at the end of the month. A practical guide to Claude in production on AWS Bedrock: model IDs, inference profiles, the four governance layers and the full cost breakdown for an internal agent.
Ox Alpha was GLM-5.3-Flash. Five days before any announcement, tokenizer fingerprinting was already pointing to Zhipu: 95 out of 95 against the GLM-5 vocabulary. Now Z.ai has confirmed it, published the weights on Hugging Face under an MIT license and revealed the architecture: 320B total with 18B active, 1M context, $0.075 per million. The 80% benchmark is still what it always was: a sample of ten tasks. On the full set, 63%.
Alibaba announced Qwen3.8-Flash-Next: 125B total with only 6B active per token, plus 51B in N-gram embeddings and a redesigned sparse attention. What Qwen has confirmed, what's still community estimate, how much memory it really needs, and why the architecture is being published ahead of Qwen 4. No official benchmark has come out so far.
One OpenAI-compatible endpoint for hundreds of models, with automatic fallback between providers. What OpenRouter is, how to call it from curl, Python, and PHP, how to control routing, provider, and cost caps, and the scenarios where this extra layer hurts more than it helps.
The internet treats the Laravel AI SDK and Prism PHP as competitors and keeps repeating that the official SDK uses Prism under the hood. The composer.json says otherwise. In this post: what each package actually does, the code side by side, the table of what only exists in one of them, the release cadence of both, and a decision guide by concrete case. No "it depends".
The Laravel AI SDK is Laravel's first-party AI package: agents as PHP classes, tool calling, structured output, streaming, embeddings with pgvector and fakes for testing without burning tokens. A practical guide from composer require to your first running agent, covering what the SDK solves, what it doesn't, and the limitations nobody tells you about.
It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.
Google ran the same play and shipped Flash before Pro. Except the $0.75/M that took over the timeline isn't a low price: the official table shows it's the 3.6 Flash price with a 50% discount through December 31, 2026, and the bill doubles on January 1. Here are both numbers, the real benchmarks (FrontierCode 43.6%, AutomationBench 30.4%), what breaks when you migrate from 3.6, and the data point the release leaves out: hallucination went up from 55.6% to 64.5%.