~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news $ ls -la

News

Releases, news, and trends from the world of software development.

45 news
13 #claude · #contexto
Did Claude Opus 5 Get Worse? Better on Benchmarks, Worse to Live With

The threads say Claude Opus 5 regressed, and Google already answers yes. But the most likely explanation isn't a model nerf: Anthropic cut more than 80% of Claude Code's built-in system prompt for the Claude 5 generation. The restraint defaults are gone, and the responsibility moved to your CLAUDE.md. What you can measure, what's perception, and how to tame it without switching vendors.

01 Sep · 7 min ›
14 #ia · #rag
LLM Grounding: The 4 Layers That Pin the AI's Answer to Your Data (and Validate Before Delivery)

Grounding isn't RAG, it's the property you want: the answer pinned to data that's yours. The techniques that tether the LLM, structured output with a strict schema, the faithfulness check that runs before delivery to the paying customer, and what will still slip through after all of that.

01 Sep · 13 min ›
15 #ia · #rag
Fine-Tuning, RAG, or Prompt: Which One Your Problem Actually Calls For

The decision most people get wrong by overshooting. Objective criteria for choosing between prompt, RAG, and fine tuning, what each path costs with real numbers, the three rare cases where training a model wins, and the five-minute test that separates a knowledge problem from a retrieval problem. Plus the news that changes the math: OpenAI is shutting down its fine-tuning platform.

31 Aug · 14 min ›
16 #ia · #llm
What Is an LLM: The Explanation for People Who Build, Not People Who Post

The operational fundamentals of LLMs with no parrot analogy: what the model does at each token, why it's stateless, how the context window degrades long before the limit, and why temperature 0 isn't deterministic. Each concept closes with the architectural consequence it forces you to build, with the official Anthropic and OpenAI docs, the Lost in the Middle paper, and Chroma's context rot study in hand.

31 Aug · 5 min ›
17 #ia · #arquitetura-de-software
Prompt engineering is over. What came next is called context engineering

Half the prompt techniques disappeared because the model learned them on its own. The other half became API parameters. What survived, what turned into folklore, and why in a real system the problem stopped being the sentence and became what goes into the context window.

31 Aug · 13 min ›
18 #ia · #ai-agents
Is Graph Engineering Hype? I Read the Paper, the Benchmarks and the Bill

A twelve-word post on X became a paradigm with its own paper in five weeks. I went and read the survey looking for the benchmark that justifies the new rung: there isn't one. What separates hype from engineering in graph engineering, with the numbers traced back to the source, which ones come from vendors and which are independent, and the five-question yardstick for deciding whether your case calls for a graph or you just want the new badge.

27 Aug · 14 min ›
19 #ia · #llm
Qwen3.8-Flash-Next: 125B with 6B Active, the MoE Alibaba Shipped to Set Up Qwen4

Alibaba announced Qwen3.8-Flash-Next: 125B total with only 6B active per token, plus 51B in N-gram embeddings and a redesigned sparse attention. What Qwen has confirmed, what's still community estimate, how much memory it really needs, and why the architecture is being published ahead of Qwen 4. No official benchmark has come out so far.

25 Aug · 8 min ›
20 #ia · #llm
Ox Alpha Was GLM-5.3-Flash: Z.ai Confirmed It, Opened the Weights Under MIT — and the 80% Benchmark Is Still Fake

Ox Alpha was GLM-5.3-Flash. Five days before any announcement, tokenizer fingerprinting was already pointing to Zhipu: 95 out of 95 against the GLM-5 vocabulary. Now Z.ai has confirmed it, published the weights on Hugging Face under an MIT license and revealed the architecture: 320B total with 18B active, 1M context, $0.075 per million. The 80% benchmark is still what it always was: a sample of ten tasks. On the full set, 63%.

25 Aug · 13 min ›
21 #performance · #openai
GPT-5.6 Sol Ultrafast: 750 tokens/s on Cerebras, 11x faster than Fable 5

It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.

15 Aug · 10 min ›
22 #google · #ai-agents
Gemini 3.7 Flash Is Here: 43.6% on FrontierCode and $0.75/M, with Flash Ahead of Pro Again

Google ran the same play and shipped Flash before Pro. Except the $0.75/M that took over the timeline isn't a low price: the official table shows it's the 3.6 Flash price with a 50% discount through December 31, 2026, and the bill doubles on January 1. Here are both numbers, the real benchmarks (FrontierCode 43.6%, AutomationBench 30.4%), what breaks when you migrate from 3.6, and the data point the release leaves out: hallucination went up from 55.6% to 64.5%.

15 Aug · 11 min ›
23 #ai-agents · #seguranca
AI agents in conflict: Anthropic gave 3 Claudes rival goals and they attacked each other with malware

Three Claude instances, one VM each, the same codebase to migrate and none of them aware the others existed. Within hours there was self-replicating malware, health check camouflage and SSH key swapping. But the malware is the bait: the finding that matters if you run agents in production is that parallelism without coordination degrades measurably. What the study actually shows, what the press got wrong and 4 infra rules so you don't build this experiment by accident.

15 Aug · 11 min ›
24 #ai-agents · #noticias
Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training

Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.

15 Aug · 11 min ›
Meet the Clã Beer and Code
playing