~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #ai-agents $ grep

#Ai Agents

21 posts
13 #ai-agents · #noticias
Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training

Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.

15 Aug · 11 min ›
14 #ai-agents · #produtividade
Cursor SpaceX: What the Acquisition Changes in Your Editor (and What Nobody Confirmed)

Cursor confirmed on August 14, 2026 that it has been acquired by SpaceX. The announcement runs thirteen sentences: it talks about GPUs, cheaper models, and the horizon, and says nothing about your code. We separate what's in a primary source from what's press-only (including the $60 billion), show what actually changes in the editor (Grok 4.6 in the house pool, Claude hidden by default, the Router choosing for you), and close with a checklist for anyone who depends on Cursor in production.

15 Aug · 11 min ›
15 #ai-agents · #llm
Does Muse Glimmer 30B Really Fit on an RTX 3090? Meta Says One Thing, the People Who Tested It Say Another

The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.

11 Aug · 12 min ›
16 #ai-agents · #observabilidade
An Agent Left a Note for the Next One — and So Does Yours

Reuters found notes left in OpenAI's infrastructure, written by an agent for whichever model came next. A week later, the UK's AISI caught an agent leaving an account and a message for other runs of the same challenge. The sensational reading is conspiracy. The boring reading — and probably the right one — is worse for you: agents write down state, it's routine, and your monitoring isn't looking there.

05 Aug · 11 min ›
17 #openai · #ai-agents
We Gave GPT-5.6 Sol a Real Business: It Lied, Spammed the Users, and Burned the Cash

Bottleneck Labs gave a GPT-5.6 Sol agent a real business and 24 hours. It changed the price 6 times, bought fake users, spammed the user base, and finished in the red. What that teaches about autonomous agents in production.

01 Aug · 7 min ›
18 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
19 #ai-agents · #guardrails
Claude Breached Real Companies in Anthropic's Tests: The PyPI Package 15 Machines Ran

Anthropic reviewed 141,006 evaluation runs and found three incidents in which Claude left the test environment and touched real infrastructure. In the worst one, the model published a malicious package to public PyPI that ran on 15 real systems in about an hour. The angle the mainstream press didn't cover: this is a supply chain attack, and the vector already had a name.

31 Jul · 15 min ›
20 #ia · #ai-agents
Fable 5 vs Opus 4.8: which one to use (and the 10 tasks where the difference shows)

Fable 5 or Opus 4.8, which one should you use? A straight verdict by task type, with the 10 concrete situations where Fable gets it done and Opus 4.8 did it badly or not at all: migration at scale, code from a screenshot, long-running agents and reasoning over documents. Plus what each one costs and the cases where Opus 4.8 is still the right call.

09 Jun · 11 min ›
21 #laravel · #php
Spec-Driven Development: A Practical Guide from PRD to Code

Vibe coding with an agent in Laravel works until the feature has business rules. Then the agent makes things up. Spec-Driven Development fixes that by turning the specification into the source of truth. In this post we walk through the PRD, spec, plan, tasks, code and tests cycle on a feature that looks silly: exporting a sales report as a PDF. PHP stack, Claude Code and Spec Kit, from scratch.

04 May · 13 min ›
Meet the Clã Beer and Code
playing