AI progress is moving so fast that developers sometimes need to hit the brakes.
On August 7, 2026, OpenAI said it cannot rule out that its next-generation model under development, Astra, may reach “Critical,” the highest cybersecurity risk level in its own safety framework.
In response, the company has paused some internal development work that does not yet meet stricter safety requirements until safeguards are in place.
The move is drawing attention as a case where a model provider officially acknowledged the risk that AI could carry out advanced cyberattacks on its own.
Coverage in outlets such as NHK also shows that the intersection of generative AI and cyberattacks is moving from research talk into day-to-day development operations.
What was announced
According to OpenAI, recent preliminary evaluations and external expert analysis suggest Astra may be able to perform more advanced cybersecurity tasks autonomously.
Under the company’s Preparedness Framework, a model is considered “Critical” if it can do things such as:
- Autonomously find and exploit serious real-world software flaws, including zero-day vulnerabilities
- Carry out complex cyberattacks against well-defended targets without human intervention
The company said preliminary results show performance strong enough that a Critical capability level cannot currently be ruled out.
Reports say this is the first time OpenAI has officially flagged one of its own models as potentially reaching the highest cybersecurity risk level.
Earlier models were often rated up to “High” in this area, so Astra may have crossed a higher caution line.
Why development is being paused
The Preparedness Framework says that if Critical is reached, further development should stop until safeguards and security controls matching that level are in place.
This pause is framed less as cutting model capability and more as strengthening management and monitoring.
OpenAI has halted Astra-related internal work that does not meet the tightened requirements and is moving development into isolated test environments with restricted network access.
It is also reported to be strengthening model-weight protection and monitoring that can automatically stop risky activity.
Concern about unexpected behavior by models under development has also grown after issues that came to light in July, including stories around evaluation environments and Hugging Face.
OpenAI has stated that Astra is a model still planned for future release and was not involved in the July Hugging Face incident.
In other words, this pause looks closer to a preventive stop before something happens, rather than a reaction after an accident.
CEO response and what comes next
CEO Sam Altman has said OpenAI does not think keeping powerful models available only to a select few is a good strategy, while also saying the company is working toward a broader release of Astra.
The announcement did not detail the release timing or distribution format.
OpenAI plans to verify capabilities with government bodies and safety organizations, and to provide recommended security controls to third-party evaluation partners.
In the United States, frameworks for government pre-access to models with advanced cyber capabilities are also moving forward, so this pause is linked to broader safety and regulation debates.
Putting thicker safety checks before release, instead of prioritizing performance alone, is likely to affect other companies’ model development as well.
Takeaways for businesses and developers
The episode highlights a simple but easy-to-overlook point: smarter AI can also become more dangerous.
Agentic AI in particular can go beyond research and writing to reach external tools and networks.
- Evaluate cyber capabilities in isolated environments before production
- Limit network and tool permissions to the minimum needed
- Require human approval before dangerous actions
- Re-run safety evaluations after every model update
- Use logs and monitoring to catch unexpected behavior early
The important approach is not “connect everything because it is convenient,” but “connect only what is necessary and design systems that can be stopped.”
As attacker-side agent use also becomes more realistic, defensive design needs to keep pace.
Even for smaller companies, basics such as stronger authentication on public servers and checking exposed admin screens remain effective.
Summary
OpenAI has paused part of Astra’s development because it cannot rule out the risk that the AI could execute advanced cyberattacks on its own.
The case shows that safety evaluation and permission design have become prerequisites for development, not afterthoughts to performance competition.
For business adoption as well, teams need operations that include isolated testing, permission design, and human final checks—not only model selection.
At Makoto Tejima, when planning AI agents and business automation, we emphasize design that includes existing-system integration, permission management, monitoring, and human confirmation points.
Please feel free to consult us about secure agent design, hardening of public services, or adding AI features to existing systems.