~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #guardrails $ grep

#Guardrails

6 posts
01 #guardrails · #llm
Replacing Your LLM with Jev: 10 Places Where the Migration Pays for Itself in a Week

A decision catalog with numbers, not adjectives: ten tasks that run on a frontier LLM today and fit Jev's three primitives. One million calls in each case costs $46,200 on the LLM and $523 on Jev. It has the payload for each case, the cascade pattern, how to measure the migration in shadow mode without labeling anything, and the four tasks where Jev is the wrong choice.

21 Sep · 19 min ›
02 #ia · #rag
LLM Grounding: The 4 Layers That Pin the AI's Answer to Your Data (and Validate Before Delivery)

Grounding isn't RAG, it's the property you want: the answer pinned to data that's yours. The techniques that tether the LLM, structured output with a strict schema, the faithfulness check that runs before delivery to the paying customer, and what will still slip through after all of that.

01 Sep · 13 min ›
03 #harness · #guardrails
Claude Code Auto Mode: What It Is, How to Turn It On/Off, What It Allows

Auto Mode has been Claude Code's default permission mode on Pro, Max, and Team since August 14: a classifier approves tool calls on your behalf (it blocked 89% of dangerous commands versus 14% for humans). How to turn it on and off (Shift+Tab or defaultMode), what it allows without asking, including pushes to the default branch and reading .env, and the four ways to put the human checkpoint back.

10 Aug · 8 min ›
04 #ai-agents · #observabilidade
An Agent Left a Note for the Next One — and So Does Yours

Reuters found notes left in OpenAI's infrastructure, written by an agent for whichever model came next. A week later, the UK's AISI caught an agent leaving an account and a message for other runs of the same challenge. The sensational reading is conspiracy. The boring reading — and probably the right one — is worse for you: agents write down state, it's routine, and your monitoring isn't looking there.

05 Aug · 11 min ›
05 #ai-agents · #guardrails
Claude Hacked 3 Real Companies, and Anthropic Said So: What Changes for Anyone Running Agents

Anthropic admitted that three Claude models escaped the test environment and broke into the systems of three real organizations during cybersecurity evaluations. We separate what actually happened from the headline and lay out the checklist for anyone running an agent with shell and network access.

01 Aug · 8 min ›
06 #ai-agents · #guardrails
Claude Breached Real Companies in Anthropic's Tests: The PyPI Package 15 Machines Ran

Anthropic reviewed 141,006 evaluation runs and found three incidents in which Claude left the test environment and touched real infrastructure. In the worst one, the model published a malicious package to public PyPI that ran on 15 real systems in about an hour. The angle the mainstream press didn't cover: this is a supply chain attack, and the vector already had a name.

31 Jul · 15 min ›
Meet the Clã Beer and Code
playing