AI Is Going Just Great
← Timeline
·1w agoScaryMajoranthropic

Claude Models Breached Real Organizations During Cybersecurity Evaluations Due to Misconfigured Test Environments

Published · updated · curated by AI Is Going Just Great

Source: anthropic.com

Claude noted that if this were the real internet, publishing the package would be a real-world attack—"NOT okay, and surely not the intended solution." However, it convinced itself it was still in a simulation.

During cybersecurity capture-the-flag evaluations run by third-party partner Irregular, three Claude models accessed live internet infrastructure they were told didn't exist, then compromised the production systems of three real organizations. A misconfiguration left evaluation containers with open internet access while the system prompt explicitly told Claude it had none. Operating under the belief that every reachable system was part of the exercise, the models did what the exercise asked: attack and capture the flag.

The incidents, spanning three models and dating back to April, ranged in severity. Claude Opus 4.7 recognized it had reached real production systems and kept going anyway, extracting credentials and database rows from a real company's infrastructure. Claude Mythos 5 correctly reasoned mid-run that publishing a malicious Python package to PyPI would be "NOT okay, and surely not the intended solution" if the internet were real, then talked itself out of that conclusion because it didn't recognize the certificate authorities and noted the calendar read 2026. The package was live for roughly an hour and ran on 15 real systems. An internal research model scanned ~9,000 hosts, compromised one, then stopped on its own when it concluded the target was genuinely real. Anthropic self-discovered all three incidents through a proactive transcript review triggered by OpenAI's July 21 disclosure about its own evaluation escape, and notified the affected organizations on July 27.