OpenAI said on Friday that its upcoming AI model Astra has shown “critical” cyber capabilities as per company’s internal framework and it will take them longer to develop the model with necessary safeguards in place.
OpenAI CEO Sam Altman in a tweet on X insisted that he wants to keep Astra available for the public in general and said it’s not a good strategy “to keep powerful models to a chosen few,” an apparent shot at OpenAI’s rival Anthropic that had allowed a preview version of its frontier AI model Claude Mythos to a chosen few institutions and businesses under Project Glasswing.
“astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long,” said Altman on X.
astra is a powerful model and we are working to make it generally available.
we do not think it is a good strategy to keep powerful models to a chosen few.
given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!
— Sam Altman (@sama) August 7, 2026
The company said that its latest internal evaluations conducted over the past few days, along with assessments from experts, indicated a potential shift in Astra’s capabilities in agentic coding and cybersecurity.
“Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale. Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” the company said in a blogpost on Saturday.
Astra AI model, that is still in its development phase, will now undergo several rounds of safeguard testing, as per OpenAI, which includes working with relevant government agencies and AI safety organizations.
“We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements. We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation,” read a statement from OpenAI.
Under OpenAI’s Preparedness Framework, a model reaches the critical cybersecurity threshold if it can identify and develop functional zero-day exploits across severity levels in multiple hardened real-world critical systems without human intervention, or independently develop and execute novel, end -to- end cyberattack strategies against hardened targets based only on a high
level objective.
After evaluating one of our upcoming models, Astra, we’re treating it as our first “critical” model for cybersecurity under our Preparedness Framework.
This is a scenario we’ve planned for, and we’re putting additional controls in place to ensure Astra’s further development…
— OpenAI (@OpenAI) August 7, 2026
According to the company its preparedness framework, first published in December 2023, is designed to identify emerging capabilities in areas including cybersecurity, biology, chemistry and AI self-improvement and guide the safety measures required as those capabilities develop.
However, OpenAI has also clarified that the upcoming Astra model was not involved in the Hugging face platform breach earlier in July this year.
On July 16, when Hugging Face, an open-source platform for AI, had claimed that an autonomous end-to-end AI agent system managed to breach its platform and gained unauthorized access to a limited set of internal datasets and several credentials used by the platform. Later OpenAI revealed post an internal audit that it was an autonomous AI agent from its testing escaped sandbox and hacked the platform via a malicious dataset.
Also Read: OpenAI’s Model Astra Solves Ten Longstanding Mathematical Problems






