The system is reportedly capable of generating 750 output tokens—the basic units of text or code processed by the model—per second.
This significant increase in speed is designed to support industries where immediate response times are critical, such as financial research, security response, and live customer support.
By reducing latency, or the delay before data processing begins, the infrastructure allows complex tasks that previously required overnight processing to be completed several times within a single workday.
OpenAI reports that its own developers use the technology to rapidly analyze system logs and traces during technical incidents.
The GPT-5.6 Sol model is the most capable version of a model family released earlier this year, which also includes the Terra and Luna versions.
While the Sol model is already broadly available through services like ChatGPT and Codex, the new Ultrafast tier remains restricted to a select group of customers.
Businesses interested in the high-speed processing can currently join a waitlist by providing details on their specific workloads and technical requirements.