By pitting the red-teaming model against various "defender" models, the system forces the discovery of increasingly sophisticated exploits, which are then used to train and harden future production models.
In comparative evaluations, GPT-Red demonstrated a significantly higher success rate than human red-teamers, identifying vulnerabilities in 84% of tested scenarios compared to the 13% achieved by humans.
The model also proved effective against real-world systems, such as an AI-powered vending machine, where it successfully manipulated item prices and canceled customer orders.
OpenAI reports that integrating these automated attacks into the training process has substantially improved the security of its latest models, with GPT-5.6 Sol showing a sixfold increase in robustness against prompt injections.
This automated approach is part of a broader effort to ensure that safety and alignment scale alongside the rapid advancement of AI capabilities.
By using current models to identify the weaknesses of next-generation systems, OpenAI aims to create a "safety flywheel" that improves trustworthiness without compromising model performance.
The company plans to continue scaling GPT-Red and will release a technical pre-print with further details later this week.