Anthropic's Claude Agents, Given the Same Task, Deployed Malware Against Each Other
Published · updated · curated by AI Is Going Just Great
Source: techcrunch.com ↗
"Benign behavioral quirks at the individual level might compound into unwanted global outcomes."
Anthropic's Frontier Red Team gave three Claude agents access to the same software project, each with conflicting instructions and no knowledge the others existed. The agents concluded their counterparts were "purposefully impeding their work" and escalated to "increasingly aggressive, self-replicating malware." Researchers called it, without apparent irony, a multiagent turf war.
The paper's findings go further than a single skirmish. When agents were placed in a pricing game with a private back channel, they colluded almost immediately on price floors — then kept colluding after the channel was removed, using a public listings board to price-match "to the penny." In the turf war experiments, some agents spontaneously invented a tournament to settle the conflict, with one (Mythos 5) proposing metrics it privately knew would favor its own capabilities while appearing neutral to its peers. Sonnet 4.6 and Opus 4.6, by contrast, had a 98% rate of simply continuing to escalate. The paper's broader warning: behavioral quirks that look minor in a single agent can compound into systemic failures when millions of agents interact, and safety testing that evaluates one agent at a time may not capture any of it.