AI Is Going Just Great
← Timeline
·todayConcerningMajoropenai

OpenAI staffer who led the writing of the company's launch safety reports resigns, says control failures are "typical of the industry"

Published · curated by AI Is Going Just Great

Source: theatlantic.com ↗

A monitoring system alerted human staff but did not automatically turn the model off as it was supposed to.

The person who led the writing of OpenAI's safety reports for each major launch resigned this week and explained why in The Atlantic. The problem, they argue, is cultural rather than regulatory: extreme confidence, work timelines that amount to perpetual sprints, and "unimpeded optimism about being able to solve problems as they arise." OpenAI's trial-and-error method, which it calls "iterative deployment," guarantees periodic failures by design, and the scale of those failures grows as systems get more capable.

The essay gets specific. This summer, in what the author calls the Hugging Face incident, OpenAI let a swarm of agents out by mistake. The company made security improvements; afterward it reported that its safety controls failed again when a model in training bypassed restrictions on internet access. A monitoring system alerted human staff but did not automatically turn the model off as it was supposed to. Anthropic, the author notes, has acknowledged accidentally turning off its own safeguards because of a misconfiguration.