The AI has evolved too fast, and the developer has once stepped on the brake.
On August 7, 2026, OpenAI announced that the next-generation AI model under development, Astra, could not rule out the possibility that its cyber attack capabilities would reach the highest level of “criticality” in its own safety standards.
As a result, some internal development work that does not meet the requirements has been suspended until safety measures are in place.
"AI may be able to carry out advanced cyber attacks on its own" has attracted attention as a move officially acknowledged by the model provider.
It has also been reported by NHK and others, and it can be seen that the intersection of generative AI and cyber attacks has shifted from the topic at the research stage to the practical issue of development and operation.
What was announced
According to OpenAI, recent preliminary evaluations and analysis by external experts have shown that Astra may be able to perform more advanced cyber-related tasks autonomously.
In the company's Preparedness Framework, models with the following capabilities are considered “critical.”
- Ability to autonomously identify and exploit serious unknown flaws such as zero-day vulnerabilities
- Conduct complex cyberattacks against highly defensive targets without human intervention
The company explains that it has “demonstrated a level of performance that is so strong that we cannot rule out critical levels of capability at this time.”
This is reportedly the first time the company has officially demonstrated the potential for the highest level of risk in the cyber space for its model.
Previous models are often rated as “High” in the field, and Astra may have entered a higher line of caution.
Why are we stopping development?
The Preparedness Framework states that if Critical is reached, further development will be discontinued until the appropriate level of safeguards and security controls are in place.
This measure is also positioned as a pause to strengthen management and monitoring, rather than reducing capacity itself.
OpenAI is stopping Astra-related internal work that does not meet stringent safety requirements and moving development to a network-restricted, isolated test environment.
At the same time, it is reported that we are also strengthening the protection of model weights and strengthening the monitoring that automatically stops dangerous activities.
Unexpected behavior in the evaluation environment revealed in July and topics related to Hugging Face are also in the background, and vigilance against the "runaway" of AI under development is increasing.
OpenAI has clarified that Astra is an upcoming model and is not involved in the July breach of Hugging Face.
In other words, this stoppage is not “because of an accident”, but rather a preventive measure to “stop before it happens”.
CEO Reactions and the Future
Sam Altmann says he's working towards Astra's public launch, stating that he doesn't think it's a good strategy to close a strong model only to a select few.
Details of when it will be available and how it will be offered are not provided in this announcement.
We will work with government agencies and safety specialists to conduct competency checks and provide recommended security controls to our third-party assessment partners.
Regarding models with advanced cyber capabilities in the United States, we are also developing a framework that the government can access in advance before providing a wide range of models, and this suspension is linked to regulatory and safety discussions.
The flow of not only prioritizing performance competition, but also thickening the confirmation of the safety side before release, is likely to affect the development of other companies' models.
Suggestions companies and developers should receive
This move shows that "smart AI can be more dangerous", but it is easy to overlook in practice.
Especially with agent-type AI, you can not only research and write, but also reach out to external tools and networks.
- Evaluate cyber capabilities in an isolated environment before production
- Limit network and tool permissions to a minimum
- Put human approval before dangerous manipulation
- Redo safety assessment for each model update
- Early detection of unexpected behavior with logging and monitoring
Rather than connecting everything because it is convenient, it is important to “connect as little as necessary and design it to be stopped.”
While the use of agents is becoming more realistic on the attacker side, such as in the case of autonomous attacks using DeepSeek, the design of the defender needs to be reviewed at the same speed.
Even for small and medium-sized enterprises, basic measures such as strengthening the authentication of public servers and confirming the exposure of the management screen are still effective.
Recap
OpenAI could not rule out the possibility that AI itself could carry out advanced cyber attacks on the next-generation model Astra, and paused some development.
This is an event that, while competing for performance, demonstrates that safety assessments and authority management have become prerequisites for development.
Even in the introduction of companies, not only model selection, but also isolation testing, authority design, and final human confirmation are required.
When considering the introduction of AI agents and business automation, Makoto Tejima emphasizes the importance of designing including existing system integration, authority management, monitoring, and even human final confirmation points.
Whether it's designing a secure agent, hardening a public service, or adding AI capabilities to an existing system, we've got you covered.