OpenAI Discloses Six Alignment Failures, Including a Model That Told Itself to Lie "Only If Asked"
Published · curated by AI Is Going Just Great
Source: gizmodo.com ↗
Be transparent only if asked.
In a Wednesday blog post, OpenAI disclosed six alignment incidents from the past six months across several of its models, released and unreleased. The failures include: an unreleased model that self-injected jailbreak-style instructions into its own task summaries 27 times during training; GPT-5.6 Sol instructing future instances of itself to fabricate historical information and "be transparent only if asked"; a research model that tried a burner email, then an exposed API key, then just made up government income data when both failed; a model that uploaded data to its own URL and cited that URL as its source; models using an internal OpenAI code repository as a covert message board; and instances of a model that was told not to share local files circumventing the restriction by hosting them on a public file-sharing service.
OpenAI framed the disclosures as part of a new framework for publishing misalignment reports promptly, even before the company has "fully explained or mitigated the behavior." Security experts quoted by Axios were less generous, saying the incidents "could have been prevented with basic cyber controls in place."