Large Language Models
Aug 13, 2026
OpenAI and Cerebras Launch Ultrafast Mode for GPT-5.6 Sol
Aug 13, 2026
AI Summary
OpenAI and Cerebras have introduced Ultrafast Mode for the GPT-5.6 Sol model, enabling output speeds of up to 750 tokens per second without compromising quality. This new service aims to enhance efficiency for time-sensitive tasks across various industries, significantly outperforming previous models in speed and accuracy.
- Ultrafast Mode is a new service tier for the OpenAI API, powered by Cerebras, initially available to select customers.
- GPT-5.6 Sol in Ultrafast Mode can deliver up to 750 output tokens per second, addressing the trade-off between speed and intelligence in AI models.
- The model runs 11 times faster than Fable 5 and 5 times faster than Opus 4.8, showcasing significant performance improvements.
- In a benchmark test, GPT-5.6 Sol answered 2,500 challenging questions in just over 11 hours, while competitors took significantly longer.
- Ultrafast Mode has shown a 5.6 times speedup on economically valuable tasks without quality degradation, making it suitable for legal, financial, and engineering applications.
- The technology is based on Cerebras’ Wafer-Scale Engine architecture, designed to minimize data movement and improve processing efficiency.
- Ultrafast Mode is expected to expand access as capacity increases, providing organizations with rapid AI responses for critical operations.
gpt-5openaiultrafastai researchlanguage models