Days after OpenAI paused training of its frontier AI models including the upcoming OpenAI Astra model over misalignment and cybersecurity risks, the company said in a statement that the model has crossed the “critical cybersecurity capability threshold”- the highest risk tier in its preparedness framework.
In a statement posted Tuesday, OpenAI claimed that Astra is the first model to cross the threshold meaning it can discover previously unknown security flaws and develop ways to exploit them without human intervention.
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible.
Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework.
We’re previewing how we evaluated…
— OpenAI (@OpenAI) September 1, 2026
However, despite the self-declared cybersecurity risks posed by the upcoming model Astra, OpenAI claimed that it has put sufficient safeguards in the model to avoid its misuse and will soon be available for public use.
Twice in August this year, OpenAI had shared that it had intentionally paused reinforcement training of its frontier AI models due to their increasing capabilities.
While Astra was not involved in the Hugging Face incident, we have incorporated our learnings from that incident into our safety approach…We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity,” said OpenAI in its statement.
OpenAI informed that they plan to make Astra available soon, “but access to its most advanced cybersecurity capabilities will be more limited and advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use.”
Meanwhile, In a post on X, CEO of OpenAI Sam Altman said that the company has slowed the development of models after Astra to give its teams more time to strengthen safety and alignment measures as AI capabilities advance.
“Astra has been done training for a while now and is a significant step forward in both capabilities and alignment. For the models after that, we have been slowing things as needed to ensure that we can do sufficient work on safety and alignment,” Altman tweeted.
Over the summer, we have been sprinting on safety priorities; it’s more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to be launching our next model soon.
There is an obvious tension…
— Sam Altman (@sama) September 1, 2026
He further mentioned that the training for Astra has been completed which represents a major step forward in both capabilities and alignment. The company is excited about the model but believes cautions is needed as AI systems become increasingly powerful.
In July, the company had claimed that its AI agent autonomously hacked Hugging Face, an open-source AI platform and community.
Meanwhile OpenAI said Astra represents a significant improvement over its previous model, GPT-5.6 Sol, particularly in vulnerability discovery, exploit development and token efficiency. In one test, the model achieved a 100 percent score on ExploitBench, a benchmark designed to assess a model’s ability to develop exploits for known vulnerabilities.
During these tests, Astra also discovered and used two previously unknown zero-day vulnerabilities as part of an exploit chain. OpenAI said it is working with the relevant maintainers to disclose the vulnerabilities.
The company mentioned that the model will be released soon, but its most advanced cybersecurity capabilities will initially be available only to a small group of testers. The access will be later expanded through Daybreak Blue, a programme intended to support defensive cybersecurity work.
“Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity,”OpenAI said.
Also Read: OpenAI Pauses Frontier RL Training for Models Including Astra Over Safety Risks





