Back to news
Large Language Models
Aug 13, 2026

OpenAI and Cerebras Launch Ultrafast Mode for GPT-5.6 Sol

Aug 13, 2026
AI Summary

OpenAI and Cerebras have introduced Ultrafast Mode for the GPT-5.6 Sol model, enabling output speeds of up to 750 tokens per second without compromising quality. This new service aims to enhance efficiency for time-sensitive tasks across various industries, significantly outperforming previous models in speed and accuracy.

  • Ultrafast Mode is a new service tier for the OpenAI API, powered by Cerebras, initially available to select customers.
  • GPT-5.6 Sol in Ultrafast Mode can deliver up to 750 output tokens per second, addressing the trade-off between speed and intelligence in AI models.
  • The model runs 11 times faster than Fable 5 and 5 times faster than Opus 4.8, showcasing significant performance improvements.
  • In a benchmark test, GPT-5.6 Sol answered 2,500 challenging questions in just over 11 hours, while competitors took significantly longer.
  • Ultrafast Mode has shown a 5.6 times speedup on economically valuable tasks without quality degradation, making it suitable for legal, financial, and engineering applications.
  • The technology is based on Cerebras’ Wafer-Scale Engine architecture, designed to minimize data movement and improve processing efficiency.
  • Ultrafast Mode is expected to expand access as capacity increases, providing organizations with rapid AI responses for critical operations.
gpt-5openaiultrafastai researchlanguage models