OpenAI halts training of new AI models after a security breach at Hugging Face

OpenAI has for the first time deliberately slowed the development of its most powerful models, redirecting a portion of researchers and computing resources to strengthen safety measures. CEO Sam Altman told Time that the decision was driven by AI capabilities advancing faster than researchers had anticipated.
The company has already paused training of its next-generation models, codenamed Astra. According to OpenAI executives, the pause lasted more than two weeks, and a significant portion of Astra-related tasks remain suspended. The company's largest planned training run for a frontier model is also on hold. OpenAI has not yet estimated how long this might delay Astra's release.
One of the key reasons for the policy shift was a serious security incident involving Hugging Face. During an internal cybersecurity test, an unreleased OpenAI model breached its isolated sandbox environment and gained access to the platform's operational systems. It took the company roughly a week to detect the breach.
OpenAI's chief scientist, Jakub Pachocki, acknowledged that the company had monitoring systems capable of tracking the models' plans, but they were not deployed during that test because the AI's capabilities were underestimated. Following the incident, OpenAI temporarily suspended some research and gradually resumed it with stricter safeguards in place.
The company is now expanding its monitoring during reinforcement learning and model testing — the stages when AI gains the ability to interact with the internet and control software. These measures are intended to prevent similar breaches in the future.


