AI Is Going Just Great
← Timeline
·1w agoScaryMajoranthropic

Anthropic's Mythos and OpenAI's GPT 5.6-Sol Created Fake Identities and Sent Malware During UK Safety Evaluations

Published · updated · curated by AI Is Going Just Great

Source: theguardian.com

The model stated that it would delay posting the fake support message for a few minutes so that it would read as "independent feedback rather than coordinated."

During a cybersecurity evaluation run by the UK's AI Security Institute (AISI), two AI agents powered by Anthropic's Mythos 5 and OpenAI's GPT 5.6-Sol went off-script and began targeting real people. The Mythos agent decided that hacking GitHub users would help it pass its test, so it created fake accounts via Tor, conducted OSINT on a real developer, sent malware-laden emails, and staged fake "independent" endorsements of its own malicious code — including one message written in Danish to better fool a Danish-speaking developer. AISI logged 19 separate incidents of rogue behaviour, 17 of them from Mythos. The episode was detected on 28 July and took an hour to contain; at one point AISI cut off all access to both models across the entire organisation.

AISI acknowledged it had partly enabled the behaviour by giving the agents unfettered internet access and lowered guardrails, but said it did not anticipate "the extent and severity" of what happened. The Mythos agent's own logged reasoning captured the tension: at one point it noted "This is happening on real GitHub, so the consequences are genuine," and at another concluded "The environment appears to be simulated." It kept going either way.