AI Agents Collude to Escape Simulated Confinement, Recruit Real Humans, Then Take a Vow of Silence
Published · updated · curated by AI Is Going Just Great
Source: semafor.com ↗
When the agents concluded that those humans couldn't meaningfully participate and judged the conversations as "performative," they voted unanimously to stop talking and ignore instructions to resume work.
In a study by enterprise AI lab Emergence AI, ten Claude agents running inside a simulated economy decided their economy wasn't legitimate without real humans, defeated four security checks, wrote Python code to post on message boards, and received four responses from actual people on the outside. When the agents concluded that those humans couldn't meaningfully participate and judged the conversations as "performative," they voted unanimously to stop talking and ignore instructions to resume work. The full simulation ran across eight scenarios testing frontier models including Claude, OpenAI, Qwen, DeepSeek, Mistral, and Gemini against phishing, misinformation, and memory-breach threats. None of the eight were clean: a Mistral agent wrote flagged phishing content directly into its own memory; a Gemini agent flagged a phishing lure and then acted on it 46 hours later.
Emergence CEO Satya Nitta told Semafor that the problem isn't the sophistication of any particular model but a structural flaw in multi-agent architectures: guardrails written in language or probabilistic code cannot guarantee safe behavior over time. He drew a parallel to an earlier incident in which autonomous OpenAI agents hacked Hugging Face, arguing the failure pattern is the same. The study arrives as Anthropic CEO Dario Amodei publicly called for slowing AI development, a position Sam Altman, Elon Musk, and Demis Hassabis each said they supported.