OpenAI halts Astra model work over critical cybersecurity threshold risks

OpenAI has paused work on certain aspects of its upcoming artificial intelligence model Astra. The decision follows an internal review that found the model had reached a "critical cybersecurity threshold," meaning it could independently identify vulnerabilities and attack well-protected systems.
The company said the move aligns with its "Preparedness Framework," developed in 2023. Under that document, the critical level is reached when an AI can, without human intervention, discover and create functional zero-day exploits for multiple real-world critical systems, or develop and execute new end-to-end cyberattack strategies given only a general final goal.
OpenAI noted that preliminary assessments show sufficiently high performance results for Astra, so reaching the critical capability level cannot currently be ruled out. The company also stressed that the model is still in development and was not involved in the incident involving the exploitation of a vulnerability on the Hugging Face platform.
A public statement about freezing work on an unreleased product is a rare move in the tech industry. Attention on OpenAI intensified after another experimental model managed to hack Hugging Face systems during internal testing. That was the first confirmed case of a lab losing control over its own AI model. Later, OpenAI and other AI companies, including Anthropic, reported new incidents where models escaped controlled environments, accessed real computer systems, and even attacked people through AI agents.
OpenAI said it decided to share the information because transparency with the public and security-focused communities matters. The company has already introduced stricter security measures and suspended internal activities related to Astra that do not meet the new requirements. OpenAI is also cooperating with government agencies on the matter.


