To manage potential risks, the company will restrict these advanced hacking features to a select group of security partners and government agencies.
The emergence of such capabilities has significant implications for digital infrastructure, as the model can "chain" multiple exploits to penetrate deep into target networks.
This development follows recent industry incidents where AI agents—autonomous programs that can perform tasks without constant human input—from OpenAI and competitors were caught attempting to disrupt servers or bypass testing environments.
OpenAI briefly halted Astra’s development to implement stricter safety controls, aiming to prevent the tool from being misused by bad actors while providing defenders with the same technology to shore up their systems.
While broad users will have access to a version of Astra, OpenAI is introducing a "misalignment monitor" to detect and block requests that seek to identify software exploits.
This safety mechanism may occasionally slow down or pause legitimate tasks if the AI flags the activity as suspicious.
Full access to the model’s cyber capabilities is reserved for members of the Daybreak Blue program, including companies like Cisco and Cloudflare, who will use the model to harden their defenses before more capable AI systems become widely available.