~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #ia $ grep

#Ia

22 posts
01 #ia · #alucinacao
What Is Jev AI: The Model That Writes Nothing (and Costs 100x Less Because of It)

TypeSafe launched Jev, the first System One model: it doesn't complete text, it takes your program's state and returns a typed decision in a single parallel pass. No parser, no retries, no broken JSON. In this post: what Jev actually is, why it can't hallucinate (and why that doesn't mean it can't be wrong), the real math on $0.042 per million tokens with free output, the three primitives with running Python code, and the 5 traps the docs themselves admit to.

21 Sep · 15 min ›
02 #ia · #observabilidade
10 AI Tools for AI Engineers in 2026 (and the Criteria for Choosing When They Change)

Tool lists rot in twelve months, yours included. That's why each of the ten comes with a swap criterion: the ten slots in the AI engineer's stack in 2026, the default for each one, and the objective signal that tells you when to rip it out. With data from the 2026 OWASP Top 10 and the Pragmatic Engineer survey.

03 Sep · 16 min ›
03 #ia · #ai-agents
When to Use an AI Agent: The 3-Minute Test

Half of the "AI projects" I reviewed in 2026 could have been solved with an if and a SELECT. The other half became agents because the team wanted to say they had AI. Here's the 3-minute test: 4 questions that decide whether your problem calls for an if, a query, a prompt, a deterministic flow or an actual agent. With 5 real cases, the verdict on each one (spoiler: 1 in 5 was an agent) and the invisible cost of overshooting.

02 Sep · 18 min ›
04 #ia · #api
JSON That Doesn't Break: LLM Structured Output Without a Defensive Parser

If your code has a try/catch wrapped around a json_decode of the model's response, you don't have a contract. You have hope. How to get truly guaranteed structured output from an LLM: constrained decoding at the provider, a schema the model can satisfy, semantic validation and retry with a fixed budget, with no infinite loop and no defensive parser.

02 Sep · 16 min ›
05 #ia · #rag
LLM Grounding: The 4 Layers That Pin the AI's Answer to Your Data (and Validate Before Delivery)

Grounding isn't RAG, it's the property you want: the answer pinned to data that's yours. The techniques that tether the LLM, structured output with a strict schema, the faithfulness check that runs before delivery to the paying customer, and what will still slip through after all of that.

01 Sep · 13 min ›
06 #ia · #produto-ia
Oracle Generative AI: Which Models OCI Has and When to Choose It Over Bedrock

The third LLM infrastructure provider already has a São Paulo region, a catalog with Grok 4.3, Gemini 2.5, Llama 4, and Cohere Command A, and bills per character instead of per token. The real Oracle Generative AI catalog, how to call the service with the OpenAI SDK, the dedicated cluster math, and the objective criteria for deciding between OCI and Bedrock.

01 Sep · 13 min ›
07 #ia · #rag
Fine-Tuning, RAG, or Prompt: Which One Your Problem Actually Calls For

The decision most people get wrong by overshooting. Objective criteria for choosing between prompt, RAG, and fine tuning, what each path costs with real numbers, the three rare cases where training a model wins, and the five-minute test that separates a knowledge problem from a retrieval problem. Plus the news that changes the math: OpenAI is shutting down its fine-tuning platform.

31 Aug · 14 min ›
08 #ia · #llm
What Is an LLM: The Explanation for People Who Build, Not People Who Post

The operational fundamentals of LLMs with no parrot analogy: what the model does at each token, why it's stateless, how the context window degrades long before the limit, and why temperature 0 isn't deterministic. Each concept closes with the architectural consequence it forces you to build, with the official Anthropic and OpenAI docs, the Lost in the Middle paper, and Chroma's context rot study in hand.

31 Aug · 5 min ›
09 #ia · #arquitetura-de-software
Prompt engineering is over. What came next is called context engineering

Half the prompt techniques disappeared because the model learned them on its own. The other half became API parameters. What survived, what turned into folklore, and why in a real system the problem stopped being the sentence and became what goes into the context window.

31 Aug · 13 min ›
10 #ia · #ai-agents
Is Graph Engineering Hype? I Read the Paper, the Benchmarks and the Bill

A twelve-word post on X became a paradigm with its own paper in five weeks. I went and read the survey looking for the benchmark that justifies the new rung: there isn't one. What separates hype from engineering in graph engineering, with the numbers traced back to the source, which ones come from vendors and which are independent, and the five-question yardstick for deciding whether your case calls for a graph or you just want the new badge.

27 Aug · 14 min ›
11 #ia · #produto-ia
AWS Bedrock: What It Is and How to Run Claude in Production with Governance (and the Bill in Reais)

AWS documents how to turn on Bedrock really well. Nobody documents the rest: the difference between CloudTrail and model invocation logging, the fact that São Paulo doesn't give you data residency, and what shows up on the bill in reais at the end of the month. A practical guide to Claude in production on AWS Bedrock: model IDs, inference profiles, the four governance layers and the full cost breakdown for an internal agent.

27 Aug · 15 min ›
12 #ia · #llm
Ox Alpha Was GLM-5.3-Flash: Z.ai Confirmed It, Opened the Weights Under MIT — and the 80% Benchmark Is Still Fake

Ox Alpha was GLM-5.3-Flash. Five days before any announcement, tokenizer fingerprinting was already pointing to Zhipu: 95 out of 95 against the GLM-5 vocabulary. Now Z.ai has confirmed it, published the weights on Hugging Face under an MIT license and revealed the architecture: 320B total with 18B active, 1M context, $0.075 per million. The 80% benchmark is still what it always was: a sample of ten tasks. On the full set, 63%.

25 Aug · 13 min ›
Meet the Clã Beer and Code
playing