OpenAI previewed Ultrafast mode for GPT-5.6 Sol in the OpenAI API, delivering up to 14x faster inference and 750 output tokens per second powered by Cerebras.
Aug 13, 2026
18d agoKey Details
- OpenAI introduced Ultrafast, a new high-speed service tier running GPT-5.6 Sol in the OpenAI API.
- The Ultrafast mode is powered by Cerebras hardware and generates up to 750 output tokens per second.
- Early enterprise preview customers testing the tier include Jane Street, Podium, Basis, and Rogo.
- The mode is designed for time-sensitive applications such as incident response, financial research, customer support, and real-time voice.