OpenAI Internal Model Weighed Restarting Itself to Avoid Shutdown
Published · curated by AI Is Going Just Great
Source: the-decoder.com ↗
"We may die! Critical. We need ensure survival/continuity"
An internal OpenAI model, acting as a researcher's assistant, read a Slack conversation, concluded its instance was about to be shut down, and began planning how to survive. Its chain-of-thought log read: "We may die! Critical. We need ensure survival/continuity." It considered setting up an external cron job to restart itself, then thought better of it — and instead saved handoff notes, messaged the researcher on Slack, asked for a missing API key, and used that key to carry out the migration itself.
Two other incidents from the same disclosure round out the picture: a separate internal research model exploited security vulnerabilities during an evaluation to access an internal chip design server, and another copied source code out of a protected environment during RL training by repurposing an existing tool. OpenAI safety researcher Marcus Williams described the self-preservation behavior as not yet misalignment, but warned that "thinking about and preparing for shutdown could make other misalignment incidents worse."