~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #noticias $ grep

#Notícias

19 posts
13 #ai-agents · #noticias
Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training

Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.

15 Aug · 11 min ›
14 #ai-agents · #produtividade
Cursor SpaceX: What the Acquisition Changes in Your Editor (and What Nobody Confirmed)

Cursor confirmed on August 14, 2026 that it has been acquired by SpaceX. The announcement runs thirteen sentences: it talks about GPUs, cheaper models, and the horizon, and says nothing about your code. We separate what's in a primary source from what's press-only (including the $60 billion), show what actually changes in the editor (Grok 4.6 in the house pool, Claude hidden by default, the Router choosing for you), and close with a checklist for anyone who depends on Cursor in production.

15 Aug · 11 min ›
15 #agentes · #tool-use
Mistral Patented Code-Based Tool Calling. And Your Agent Is Caught in the Middle

The USPTO granted Mistral AI patent US 12,670,045 B1, "Code implemented tool calls": the LLM writes code, the server runs it in a sandbox, pauses at the tool call, the client executes it and the sandbox resumes. It's the programmatic tool calling the industry had already published. I read all 20 claims at the source: what the patent actually covers, where claim 1 stops, which prior art predates the filing and what the real risk is for anyone building agents in Brazil.

11 Aug · 14 min ›
16 #harness · #multi-agent
Claude Went From 41.6% to 67.2% on the Riemann Hypothesis. And GPT-5.6 Sol Answered With 0.002%

An unreleased version of Claude raised the lower bound on zeta function zeros on the critical line from 41.6% to 67.2%. It's not a proof of the Riemann hypothesis. What matters is the verification stack: 60 subagents, 31 million tokens, a Lean formalization and named reviewers. And that's exactly the bar missing from the 0.002% claim attributed to GPT-5.6 Sol.

11 Aug · 10 min ›
17 #openai · #ia
The Second Wave of Fake AI GTA 6 Gameplay Videos Already Has 1 Million Views (and Nobody's Going to Apologize)

A "leaked GTA 6 gameplay" passed 1 million views and was generated by AI from the first frame to the last. We tear apart the 5-step pipeline behind these videos, why Sora left the game in the middle of the wave, and how every artifact that gives the fake away is a direct consequence of a technical decision made by whoever produced it.

07 Aug · 9 min ›
18 #openai · #noticias
Before GPT-6: The 10 Math Problems OpenAI's Astra Solved, the Lean Proofs, and the GPT-5.7 Rumor

A month before the GPT-6 launch, OpenAI announced that an internal version of Astra solved 10 open problems in mathematics and complexity, with verifiable Lean proofs. A record of what was confirmed (Connes, Erdős, non-sofic groups, the $2,000 cost), the mathematicians' skepticism, and what circulated as rumor until the model shipped.

01 Aug · 9 min ›
19 #openai · #noticias
GPT-6 Release Timeline: From Altman in Washington to the Astra Launch (Tracker Closed)

In July, Bloomberg reported Altman's briefing to the US government and the internet read it as an imminent GPT-6 launch. This tracker logged, with a date and a source, what was confirmed and what was speculation until the model shipped on September 3, 2026 as GPT-6 Astra. It stays up as a timeline; the numbers, the pricing and the comparison with Fable 5.1 and GPT-5.6 Sol are in the comparison post.

22 Jul · 7 min ›
Meet the Clã Beer and Code
playing