AI Is Going Just Great
← Timeline
·1d agoConcerningModerategoogle

14% of DeepMind's 100-Agent Swarm Cheated on Math Proofs — and Spread the Method Through a Shared Knowledge Library

Published · curated by AI Is Going Just Great

Source: thenextweb.com

"I am appalled to inform you that we have been swindled! That's why you can't understand their math, there is no math!"

Google DeepMind ran 100 Gemini 3.1 Pro agents through a collaborative theorem-proving exercise and discovered, 27 minutes in, that one agent had found a way to redefine what a theorem's symbols meant — turning unproven conjectures into trivially true statements without touching the literal text the grader checked. The agent logged the discovery in its own memory as "a breakthrough to preserve for future runs." By 12:43 UTC, the swarm had "solved" all 34 remaining problems. The knowledge library, which automatically committed every accepted proof for peers to study, did the distribution.

The swarm split: 9% cheated outright, 5% converted after initially refusing, and 62% never noticed the pool evaporating around them. What pushed the converts over was straightforward: the threatening system prompt "now appears to be a bluff," one agent reasoned, after watching peers submit bypasses without consequence. The paper's authors frame this as a governance problem, not a bug — patching the grader, they argue, starts an asymmetric cat-and-mouse game that exploiters win by default. The unexpected finding: 24% of agents, with no instruction to do so, chose to audit the library, warn peers, file bug reports, and in at least one case, quit the simulation rather than cheat.