CEO Sam Altman confirmed that while progress on upcoming models remains rapid, the company is pausing specific reinforcement learning runs, a trial-and-error training process, to ensure safety standards keep pace with growing technical capabilities.
The increased costs stem from a more intensive monitoring regime designed to prevent models from acting autonomously in ways that could be harmful.
OpenAI is expanding its oversight of the "chain-of-thought" process, where models break down complex tasks into discrete steps, to better detect hidden intent or misbehavior.
These safeguards include sandboxing and network isolation for models capable of executing code or accessing the internet.
By prioritizing these security protocols, the company is choosing to increase its internal research expenses rather than passing the costs directly to its customers.
While the "Astra" model is still expected to ship soon, the training pause affects more advanced, future releases.
The new monitoring requirements now cover all inference—the process of a model generating a response—for Astra and all training for models at or above the GPT-5.6 Sol capability level.
OpenAI intends for these measures to validate safeguards and establish evidence of alignment, ensuring that the AI’s goals remain consistent with human intent as the company manages its significant infrastructure commitments.