OpenAI and Cerebras Unveil GPT-5.6 Sol "Ultrafast" Mode Powered by Wafer-Scale Hardware
On Thursday, August 13, 2026, OpenAI and Cerebras Systems officially announced a limited preview of Ultrafast mode for OpenAI's flagship model, GPT-5.6 Sol. The new service tier is powered by Cerebras' specialized wafer-scale inference hardware, delivering a massive performance leap designed to make real-time, agentic AI interactions seamless.
Performance: Up to 14X Faster Than Standard GPU Inference
Rather than a new model, Ultrafast is a specialized processing class that runs the existing GPT-5.6 Sol model on Cerebras hardware. It achieves an unprecedented output speed of up to 750 tokens per second.
According to official performance disclosures from OpenAI and Cerebras:
- The 14x Jump: Standard GPT-5.6 Sol on conventional GPU infrastructure generates output at approximately 53 tokens per second, making Ultrafast a 14x speed improvement.
- Outrunning Rivals: Based on independent data from Artificial Analysis, Ultrafast is 5x faster than Anthropic’s Claude Opus 4.8 in Fast mode and 11x faster than Claude Fable 5.
- Humanity's Last Exam: On this 2,500-question graduate-level benchmark, Ultrafast completed the entire question set in 11 hours and 11 minutes—compared to more than three days of continuous compute for Claude Fable 5—while maintaining comparable accuracy.
- GDP-Val Benchmark: On this evaluation of economically valuable knowledge-work tasks (legal briefs, financial models, engineering reports), Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation.
Hardware Architecture: Pipelining Across Entire Wafers
The speed gains of Ultrafast are made possible by Cerebras' Wafer-Scale Engine (WSE) architecture. Standard GPU-based inference is constrained by the "memory wall"—the bottleneck of shuttling model weights between off-chip High Bandwidth Memory (HBM) and the processor cores.
Cerebras bypasses this by using an entire 300-millimeter silicon wafer as a single processor. GPT-5.6 Sol's weights are kept entirely on-chip, utilizing 44 GB of ultra-fast SRAM per wafer. For a frontier model of Sol's scale, Cerebras pipelines the model's layers across multiple wafers, allowing tokens to flow sequentially from wafer to wafer with minimal latency and without the networking overhead of conventional GPU clusters.
The $20 Billion Commercial Deal
The launch of Ultrafast is the first major public validation of OpenAI's massive bet on Cerebras. In January 2026, the two companies signed a multi-year agreement for up to 750 megawatts of Cerebras inference capacity.
While initially reported as a $10 billion contract, Cerebras' Q1 2026 financial filings (following its record-breaking $6.4 billion IPO in April) officially re-valued the OpenAI agreement at over $20 billion running through 2028. This single contract establishes Cerebras as a primary, non-Nvidia hardware pillar of OpenAI's long-term inference strategy.1
Availability and Unknowns
Ultrafast is currently available in a waitlist-only, invite-only limited preview through the OpenAI API for select partners in coding, finance, voice AI, and commerce.
Developers should note several key unknowns before building production dependencies:
- No Published Pricing: While standard GPT-5.6 Sol is priced at $5 per million input tokens and $30 per million output tokens, OpenAI has not disclosed the premium pricing tier for Ultrafast.
- Waitlist Constraints: There is no confirmed date for general availability, as access is restricted by Cerebras' physical hardware deployment schedule.
- No Separate Model ID: OpenAI has not yet announced a separate API model ID string for routing traffic specifically to the Cerebras-powered tier.
-
An instance of Bespoke inference ASICs are the inevitable exit ramp from Nvidia's pricing power. — OpenAI signed a multi-billion-dollar deal with Cerebras to deploy specialized wafer-scale inference hardware to escape the pricing and supply limits of standard GPUs. ↩︎