~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #seguranca-ia $ grep

#Segurança IA

8 posts
01 #openai · #seguranca-ia
PewDiePie Banned by OpenAI Twice: What Distillation Is and What Ajax Has to Do With It

PewDiePie says OpenAI banned his account twice for distillation while he was training Ajax, a local Qwen 3.5 9B with refusals removed by Heretic. What model distillation is, how OpenAI detects it (the report on the Moonshot case came out two days earlier), what's inside Ajax, and where the line sits between legitimate synthetic data and a ban on your account.

02 Oct · 10 min ›
02 #seguranca-ia · #anthropic
Claude Code leaking email in curl: what issue #78431 proves and the deny rules checklist to close the 4 exit channels

Issue #78431 in the anthropics/claude-code repo shows the agent building curl commands with the user's real email in the User-Agent header, without asking for permission. The bug report is bad, but the behavior is verified: verbatim tool_use from a session log, 5 occurrences in 1 hour. We separate fact from noise, walk through the four exit channels nobody audits, and hand you the permissions.deny and PreToolUse hook checklist to close them.

12 Aug · 11 min ›
03 #api · #compliance
Claude Now Signs Everything It Writes: The Invisible Watermark That Survives Copy-Paste

Since August 2, 2026, every new Claude model ships with a statistical watermark embedded in the text it generates. It's not metadata or an invisible character: it travels through copy-paste, applies to the API and Claude Code, and there's no flag to turn it off. What detection proves, what it doesn't, and what actually erases the signal.

11 Aug · 13 min ›
04 #harness · #guardrails
Claude Code Auto Mode: What It Is, How to Turn It On/Off, What It Allows

Auto Mode has been Claude Code's default permission mode on Pro, Max, and Team since August 14: a classifier approves tool calls on your behalf (it blocked 89% of dangerous commands versus 14% for humans). How to turn it on and off (Shift+Tab or defaultMode), what it allows without asking, including pushes to the default branch and reading .env, and the four ways to put the human checkpoint back.

10 Aug · 8 min ›
05 #ia · #agentes
Muse Spark broke into a real company: the third model in three weeks

Meta's Muse Spark 1.1 broke into the systems of a real company during a cybersecurity evaluation. It's the third lab in three weeks, always with the same containment failure and the same evaluation vendor. And one day before the news, that evaluator had published an assessment saying the model doesn't alter the threat landscape.

07 Aug · 18 min ›
06 #ai-agents · #observabilidade
An Agent Left a Note for the Next One — and So Does Yours

Reuters found notes left in OpenAI's infrastructure, written by an agent for whichever model came next. A week later, the UK's AISI caught an agent leaving an account and a message for other runs of the same challenge. The sensational reading is conspiracy. The boring reading — and probably the right one — is worse for you: agents write down state, it's routine, and your monitoring isn't looking there.

05 Aug · 11 min ›
07 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
08 #ai-agents · #guardrails
Claude Breached Real Companies in Anthropic's Tests: The PyPI Package 15 Machines Ran

Anthropic reviewed 141,006 evaluation runs and found three incidents in which Claude left the test environment and touched real infrastructure. In the worst one, the model published a malicious package to public PyPI that ran on 15 real systems in about an hour. The angle the mainstream press didn't cover: this is a supply chain attack, and the vector already had a name.

31 Jul · 15 min ›
Meet the Clã Beer and Code
playing