Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras wafer-scale hardware. The tier runs OpenAIs GPT-5.6 Sol at up to 750 output tokens per second with no quality compromise, targeting time-sensitive, mission-critical workloads. Access starts with a select group of customers and expands over time.
The speedup is dramatic in independent comparisons. According to output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast runs about 11x faster than Claude Fable 5 and 5x faster than Opus 4.8 on Fast mode. In Cerebras own test on Humanitys Last Exam, Ultrafast answered all 2,500 questions in 11 hours and 11 minutes, while Claude Fable 5 needed 78 hours and 27 minutes for the same task, achieving comparable accuracy roughly 7x faster.
On GDP-Val, a benchmark for economically valuable knowledge work, Ultrafast delivered a 5.6x end-to-end speedup. For founders, that means frontier reasoning can sit on the critical path of products where latency matters: legal briefs, financial models, engineering reports, and agent loops. Cerebras argues fast frontier inference is fundamentally a data movement problem, and its Wafer-Scale Engine architecture is built to solve exactly that.
