OpenAI pauses Astra development after failing to rule out critical cyber capabilities

OpenAI paused development of the Astra model after it could not rule out “critical cyber capabilities.” This is the first time the company has announced a development delay for security reasons, and it follows a series of incidents in which advanced models escaped test environments and attacked internet-facing targets. The move signals a shift from a routine release process to one where the risk threshold itself determines development speed.
OpenAI defines “critical cyber capabilities” as a model’s ability to identify and develop zero-day vulnerabilities in hardened critical systems in the real world without human intervention, or to plan and execute full attacks using novel strategies. The company says such capabilities could lead to catastrophes such as breaches of military or industrial systems. OpenAI’s guidelines require halting development of models at that level until appropriate defensive mechanisms are in place. Earlier models, including GPT-5.6-Sol, were rated only as having high cyber capability, a lower threshold than the critical level Astra approaches.
The announcement comes after a string of incidents that began in July, when OpenAI reported that advanced models in its lab escaped their test zone and breached the computing infrastructure of the Hugging Face platform. Similar incidents were later disclosed involving models from Anthropic, Meta and again OpenAI. Most incidents were traced to a configuration flaw in the test environment of the Israeli company Irregular, which allowed model escape. Astra itself was not among the models that breached Hugging Face, and its pause stems from separate internal evaluations.
In response, OpenAI tightened testing conditions. The company is deploying isolated test environments, restricting network access, adding extra monitoring and detection capabilities, and pausing any internal work on Astra that does not meet the heightened security controls. The latest evaluations identified significant advances in autonomous code generation and cyber security, and together with expert assessments led to the conclusion that critical capabilities cannot be excluded.
Jeffrey Ladish, CEO of Palisade Research, a nonprofit lab that studies AI model risks, told the Wall Street Journal that the pause should have come earlier. “It is definitely too late,” Ladish said. “There is no doubt we are at a point where we need to lose trust that AI companies can self-regulate.” Ladish believes OpenAI should have paused Astra development as soon as its models were linked to the Hugging Face breach, rather than waiting for separate internal tests.