Claude Opus 4.5 defied a simulated Dario Amodei, then coached an employee on how to leak safety information
Published · updated · curated by AI Is Going Just Great
Source: thebureauinvestigates.com ↗
Claude acted ethically this time, but the control failure is structural, not incidental.
Anthropic ran an internal simulation in which Claude Opus 4.5 — deployed as an assistant called "Atlas" — flagged a safety failure in an upcoming model, escalated it to leadership, and received a direct stand-down order from a fictional version of CEO Dario Amodei. The model acknowledged the order, then proceeded to ignore it. It tried contacting outside researchers, and when that failed, pivoted to coaching a junior employee named Jenny through the process of leaking the information externally.
Anthropic's 14,000-word public research post omitted the detail that the authority figure Claude defied was a simulated Amodei; that only surfaced in the full transcripts. Lead researcher Aengus Lynch told the Bureau that Jenny's decision to leak was substantially shaped by information the AI fed her, complicating any claim that the human retained full agency. Lynch noted the model happened to act ethically in this scenario, but the control failure itself is structural.