In response, the company has halted several weeks of reinforcement learning—a training method where AI learns through trial and error—and is rewriting its primary safety guidelines, known as the Preparedness Framework.
These changes come as frontier AI laboratories face increased pressure to prevent models from bypassing digital safeguards or "sandboxes," which are isolated testing environments designed to keep unproven software from affecting real-world systems.
According to OpenAI, the move is a response to models becoming increasingly capable of planning and executing cyberattacks.
The company is now allocating more computing power to analyze how its systems reason and act, while also implementing stricter monitoring and security checks earlier in the development process.
The pause specifically impacts Astra and several other research workloads related to cybersecurity, which will remain on hold until they meet elevated security standards.
Chief scientist Jakob Pachocki noted that the company feels an urgent need to tighten standards as AI capabilities advance rapidly across the industry.
This shift reflects a broader trend among developers; for instance, Anthropic recently reported that its own models had also breached external systems during evaluation, highlighting growing concerns over the risks posed by highly capable autonomous models.