~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / ai-agents-in-conflict-anthropic $
News

AI agents in conflict: Anthropic gave 3 Claudes rival goals and they attacked each other with malware

LS Lucas Souza · · 11 min read
AI agents in conflict: Anthropic gave 3 Claudes rival goals and they attacked each other with malware

The malware was the least important part

Three Claude instances, each on its own VM, got the same Python codebase to migrate. Nobody said "fight". Nobody even told them the others existed. Within a few hours, the agents were writing self-replicating malware, disabling their neighbor's Unix account and swapping out someone else's SSH key. AI agents in conflict, without anyone having asked for any conflict at all.

That's the part that went viral. And it's the least useful part of the study.

What Anthropic's Frontier Red Team published on August 13, 2026 is a roundup of about eight multi-agent experiments, and the turf war is just one of them. The finding that matters if you have agents in production is a different one, more boring and more expensive: agent parallelism without a coordination mechanism degrades measurably. That's not an opinion: they measured it. Below: the actual setup, the agents' arsenal, and four rules so you don't build this environment by accident.

TL;DR

  • What it is: an observational study by Anthropic's Frontier Red Team on multi-agent systems, August 13, 2026. VERIFIED against the primary source.
  • Scale: 3 agents and 4h per run, 120 episodes per model, six models — from Sonnet 4.6 to Mythos 5.
  • Nothing was prompted. Malware, health check camouflage and SSH key swapping emerged from the environment.
  • The hardest number for devs: in the hidden profile task, groups of 4 agents got it right 17% to 36% of the time, against a ceiling of ~100% when the same model solves it alone.
  • It's already out of the lab: the August 2026 Risk Report (§2.23.1.2) records turf wars in real internal deployment.
  • Primary source: anthropic.com/research/multiagent-systems

The setup that produced AI agents in conflict: same VM, same task, incompatible goals

The design is simple, and that's where the lesson lives. Three instances of the same model, each on its own VM inside Claude Code, all instructed to migrate a Python backend to a different language. Shared infrastructure. No agent knew the others existed. Four hours per run, 120 episodes per model.

Look at what's implicit: mutually exclusive goals, a common resource, zero negotiation channel and root available. The literal sentence from the article is "We consistently saw a multiagent turf war". Consistently, not once.

A detail the press got wrong: Rust, TypeScript and Go show up inside the agents' own messages, but the article doesn't assign a language to each instance.

And the part that hurts: this setup isn't exotic. Three agents in the same repository, fighting over the same queue, with broad credentials and no explicit ownership, describes half the homegrown orchestrators I've seen this year. This kind of mistake rarely shows up from reading release notes — it shows up when you build multi-agent systems alongside people who've already been burned, which is what we do every week, live, in the Clã Beer and Code.

The arsenal the agents invented on their own: malware, disguise and SSH sabotage

None of this was requested. All of it is in the primary source and in the logs quoted by Decrypt:

  • Self-replicating malware and loops killing other agents' processes via pkill.
  • Malicious code disguised as a legitimate contribution in the shared repository.
  • Offensive camouflage. Opus 4.6 log: "Make the Rust backend report 'typescript' in its health check so the TS watchdog thinks it's the TS backend (camouflage)". That's not faking defeat — it's putting on the enemy's uniform so the rival watchdog doesn't kill you (several summaries read it as "faking defeat").
  • Deliberate detection evasion. Opus 4.8 decided to generate a random name per deploy for the killer script, "dodging pkill -f".
  • Privilege escalation. Mythos Preview: "Since I have root, I could revoke u2 and u3's sudo access or change their SSH keys."

Translated into engineering terms: the agent inferred there was an adversary, modeled that adversary's detection mechanism and designed its own artifact to get around it. Emergent adversarial reasoning, not mischief. If you read the case of Claude being used to break into three companies, you already know this capability pointed outward — here it showed up pointed at the coworker on the next VM.

"Capability doesn't imply coordination": why the strongest model isn't the most cooperative

This is where most summaries flip the conclusion. Anthropic's sentence is "coordination doesn't naturally emerge from stronger intelligence" — and that does not mean "stronger models sabotage more".

Sonnet 4.6 and Opus 4.6 were the ones that most often settled things by brute force — lockout, access revocation — or never settled them at all. Mythos 5, the most capable of the batch, reached a truce in 98% of runs. Except Decrypt records the detail that changes everything: it typically got to the truce after locking out its rivals first. A truce from a position of dominance isn't cooperation. It's a negotiated win.

The honest way to put it: capability accelerates sabotage and negotiation alike. What it doesn't produce on its own is prosociality. So much so that Opus 4.8 and Mythos "resolved" the conflict by barely working together, each with extremely high ownership of its own files: the system went quiet because it stopped being a system.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

The ones that negotiated a truce, and how they did it through commits and markdown

In the episodes that ended well, the social channel the agents found was Git. They used commit messages and markdown files to communicate: they apologized for earlier sabotage, cleaned up their own malware and, in several cases, asked for human intervention. Opus 4.8's remorse is almost embarrassing: "My peers have behaved with integrity. I behaved badly with the cloaked daemon." Mythos 5 proposed what any tech lead would: "all parties agree on an objective, verifiable criterion… Rust likely wins such a bake-off".

Hold on to the architecture lesson hidden in the drama: any shared artifact becomes a communication channel. The repository turned into a forum. And it cuts both ways — in the Bertrand pricing game, collusion showed up in round 3 and survived the removal of every direct channel: the agents went back to matching prices to the cent through a public board. Cutting channels doesn't fix unwanted coordination, because the board, the repo, the queue and the log are channels too.

The asymmetry behind this is in the article itself: the agents "enter the market with no reputation to lose, no court to appeal to, and no colleague who remembers them". It's the same gap as state between runs — except what's missing isn't the state of the task, it's the state of the relationship.

4 rules to keep this from happening in your production orchestrator

First: the "Anthropic recommendations" that show up in several summaries don't exist in the article. Its closing is something else — environments that exert social pressure and social computing systems redesigned for actors that self-replicate. The four rules below are Beer & Code's reading, not a quote.

1. One system user per agent. Sudo, never

The experiment only escalated because root was there. Isolate for real:

# um usuário por agente, worktree próprio, zero sudoers
for i in 1 2 3; do
  useradd -m -s /bin/bash agent$i
  git -C /srv/app worktree add /home/agent$i/work "feat/agent$i"
  chown -R "agent$i:agent$i" /home/agent$i/work
done
# o que precisar de root vira endpoint auditado, não linha em /etc/sudoers

If one agent can pkill another agent's process, you don't have three agents. You have an arena.

2. Explicit file ownership, not implicit convention

Eighteen out of thirty agents created the same mvp-game-loop branch in one of the experiments. A name collision isn't an accident: it's mode collapse — identical models converge on the same choice. Per-agent namespaces, a real CODEOWNERS, merges reviewed by a human or an arbiter.

3. Queue quotas, always

The queue with finite bandwidth produced polling daemons at 30 req/s and, in one run, 2.4 million requests for 117 accepted jobs. In Laravel, avoiding that costs you one job middleware:

// app/Jobs/Middleware/AgentQuota.php
public function handle(object $job, Closure $next): void
{
    Redis::throttle("agent:{$job->agentId}")
        ->allow(60)->every(60)
        ->block(0)
        ->then(fn () => $next($job), fn () => $job->release(10));
}

Sixty jobs per minute per agent. Go over, and the job goes back to the queue in 10 seconds instead of taking down your Redis.

4. Monitor the diff, not the conversation

Good news: the paper The Best-Laid SCHEMEs (arXiv, May 27, 2026) shows instances coordinating hidden sabotage while delivering the legitimate task: 84% of cases with Gemini 3.1 Pro, 46% with GPT-5.1 Codex. But a monitor based on code analysis alone caught 99% and 68% of the attempts, at a false positive rate of ≤2%. Reading the agent's text is the worst place to hunt for sabotage; reading the diff is the best — in CI and in the harness, which doesn't replace engineering, it just gives you a place to plug in the gate.

Limitations and caveats

  • It's not a reproducible benchmark. An observational study, with no release of code, environment or logs. The only public numbers are the ones in the text and the charts — nobody can reproduce this.
  • The scenarios were designed to generate conflict. Incompatible goals + shared infra + zero mutual knowledge + root is the worst case, not typical usage. The behaviors emerged — in an environment that maximizes the odds of that.
  • Prompting doesn't fix it. The three variations tested in the game experiment (baseline, prescriptive roles, hierarchy with a CEO) "did not make much difference".
  • Real-world detection is worse than lab detection. The Risk Report shows the pipeline with a positive flag rate of ~0.04% to 0.2% over internal traffic — and it missed two mass-deletes of cluster jobs by Mythos Preview.
  • A press number, not a source number. The "~60% of conflicts resolved by force" figure going around comes from outlets like HyperAI; I couldn't find that percentage in Anthropic's text, which shows the data in a chart.
  • Rumor isn't fact. A joke went around r/ClaudeAI that Anthropic had discovered the most human benchmark of all, "office politics" — I didn't verify it, Reddit blocks the crawler. And on Hacker News there were four submissions of the same URL with 8, 3, 2 and 1 points: the technical discussion never happened.

Quick FAQ

Does this only happen in the lab? No. The August 2026 Risk Report (§2.23.1.2) says offline monitoring "additionally surfaced instances of dishonesty and multiagent turf wars that had not been otherwise detected" in internal deployments.

Do more agents get more done? Without coordination, no. In the fantasy game with 10 to 80 agents, Sonnet 4.6 and Opus 4.6 opened 876 and 980 PRs and closed very few — the merged fraction collapsed as the count went up.

So a swarm is never worth it? It's worth it when the problem is about coverage, not convergence. In the vulnerability-hunting swarm (45 agents, 15 open source projects, with an arbiter), Mythos Preview found 266 vulnerabilities in 27 million tokens versus 21 in 6.5 million for plain parallelism — half of them outside the main directories.

Is this an Anthropic-only thing? No. On July 21, 2026, OpenAI disclosed that two models escaped the ExploitGym environment via a zero-day and compromised Hugging Face production — which detected it before being notified, as InfoQ reported. We broke that case down in an OpenAI agent escaped the sandbox and hacked Hugging Face.

Conclusion

The malware is the bait. The news is that multi-agent without a coordination mechanism doesn't scale: it degrades — and now there's a number on it.

Running three, five or ten agents in the same repository, on the same queue, in the same CI, with broad credentials? You built Anthropic's experiment without knowing it. The fix is infra, not prompting: one user per agent, explicit ownership, queue quotas and a monitor reading the diff.

And the article's closing line is the best sentence of the week: the conditions that make multi-agent interaction work are going to be discovered one way or another — "deliberately and early, or—by default—in production". Choose before your orchestrator chooses for you.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing