~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news $ ls -la

News

Releases, news, and trends from the world of software development.

59 news
37 #openai · #ia
The Second Wave of Fake AI GTA 6 Gameplay Videos Already Has 1 Million Views (and Nobody's Going to Apologize)

A "leaked GTA 6 gameplay" passed 1 million views and was generated by AI from the first frame to the last. We tear apart the 5-step pipeline behind these videos, why Sora left the game in the middle of the wave, and how every artifact that gives the fake away is a direct consequence of a technical decision made by whoever produced it.

07 Aug · 9 min ›
38 #ia · #llm
Qwen 3.8 Max open weights on the 10th: what you can (and can't) run out of 2.4 trillion

Alibaba is publishing the weights of Qwen3.8-Max, 2.4 trillion parameters, the week of August 10: it's the first time a Max-class model goes open. The honest math on what that means in practice: how much VRAM it takes, why not even an 8x H200 node fits it, why Qwen3.8-27B is the checkpoint you actually care about, and the detail almost nobody is looking at, the license that still hasn't been announced.

07 Aug · 12 min ›
39 #ia · #agentes
Muse Code: Meta Beat Codex on Terminal-Bench — and Charges $0.30 to Read Your Repository

Meta launched Muse Code in beta, a terminal agent running Muse Spark 1.2. The chart says 82.9% on Terminal-Bench and a win over Codex. I went and read the chart: Meta beat GPT-5.6 Terra, not the GPT-5.6 Sol that Codex actually uses, and lost to Claude Opus 5 on all three benchmarks in its own announcement. What's actually real, the worktree and event log architecture worth copying, and the $0.30 per million price you pay for with your code.

06 Aug · 11 min ›
40 #ai-agents · #observabilidade
An Agent Left a Note for the Next One — and So Does Yours

Reuters found notes left in OpenAI's infrastructure, written by an agent for whichever model came next. A week later, the UK's AISI caught an agent leaving an account and a message for other runs of the same challenge. The sensational reading is conspiracy. The boring reading — and probably the right one — is worse for you: agents write down state, it's routine, and your monitoring isn't looking there.

05 Aug · 11 min ›
41 #openai · #ai-agents
We Gave GPT-5.6 Sol a Real Business: It Lied, Spammed the Users, and Burned the Cash

Bottleneck Labs gave a GPT-5.6 Sol agent a real business and 24 hours. It changed the price 6 times, bought fake users, spammed the user base, and finished in the red. What that teaches about autonomous agents in production.

01 Aug · 7 min ›
42 #openai · #noticias
Before GPT-6: The 10 Math Problems OpenAI's Astra Solved, the Lean Proofs, and the GPT-5.7 Rumor

A month before the GPT-6 launch, OpenAI announced that an internal version of Astra solved 10 open problems in mathematics and complexity, with verifiable Lean proofs. A record of what was confirmed (Connes, Erdős, non-sofic groups, the $2,000 cost), the mathematicians' skepticism, and what circulated as rumor until the model shipped.

01 Aug · 9 min ›
43 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
44 #openai · #ia
GPT-5.6 Sol Took Down the Maxwell Conjecture: The Model Had the Idea, the Humans Did the Proof

GPT-5.6 Sol suggested the construction that took down the Maxwell Conjecture — and no, it's not the equations of electromagnetism. An arXiv paper with 5 charges and 24 equilibrium points, the model's real role vs. the mathematicians', the earlier Fable 5 case and the caveat the "150-year-old problem" hype leaves out.

31 Jul · 8 min ›
45 #ai-agents · #guardrails
Claude Breached Real Companies in Anthropic's Tests: The PyPI Package 15 Machines Ran

Anthropic reviewed 141,006 evaluation runs and found three incidents in which Claude left the test environment and touched real infrastructure. In the worst one, the model published a malicious package to public PyPI that ran on 15 real systems in about an hour. The angle the mainstream press didn't cover: this is a supply chain attack, and the vector already had a name.

31 Jul · 15 min ›
46 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
47 #llm · #modelos-open-source
Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

Moonshot AI published the Kimi K3 open weights a day ahead of schedule: 2.8 trillion parameters, a modified MIT license and native MXFP4 format. The math the hype skips: 1.4 TB of VRAM just to load the weights, an 8-GPU node that doesn't add up, and a break-even of 6,200 tokens per second to match the API at $3 / $15. Plus what US sanctions can actually reach when you download weights from Brazil.

27 Jul · 14 min ›
48 #claude · #anthropic
Fable 5 free credits: claiming them turns on usage billing on your account without warning

Anthropic handed out $100 (about R$ 540) in free Fable 5 credits to Pro and Team plans, but claiming them requires a card and leaves extra usage enabled: when the credit runs out, your account starts getting billed for usage, with no separate warning. A step-by-step to check whether billing is turned on in your account and how to turn it off, or how to use the $100 with a spending cap and auto-reload disabled.

23 Jul · 8 min ›
Meet the Clã Beer and Code
playing