According to OpenAI VP of research Amelia Glaese, the model has demonstrated the ability to independently discover and exploit previously unknown security flaws in well-protected systems without human guidance.
The decision to limit access highlights the difficulty of balancing security innovation with the risk of misuse.
Astra represents a leap in AI infrastructure, as it successfully identified and chained together two "zero-day" vulnerabilities—security holes unknown to the software's creators—during internal testing.
While these capabilities could help organizations find and patch weaknesses, OpenAI warns that without strict safeguards, the model could significantly increase the effectiveness of malicious attackers or take unauthorized actions on its own.
As part of these new safety measures, OpenAI warned that Astra’s safeguards might mistakenly flag legitimate activity as a security threat.
This could lead to the model pausing or stopping tasks, particularly long-running agent tasks—AI processes that operate autonomously to complete complex goals.
While individual users may be asked to review flagged actions, tasks performed through the company's API (the interface allowing other software to connect to the model) will simply stop, potentially affecting businesses and developers using the technology for non-security work.