❯ OpenAI Previews Ultrafast Service Layer, GPT-5.6 Sol Up to 14x Faster
SHIFTOpenAI is previewing a new API service layer, Ultrafast: powered by Cerebras wafer-scale chips, it runs GPT-5.6 Sol at output speeds up to 750 tokens/second, up to 14x faster than standard processing, with intelligence on par with the standard version. It is currently in limited release to a select group of customers, with gradual rollout as capacity expands.
HARDWAREThe speed comes from Cerebras’s wafer-scale engine architecture: each wafer-sized chip carries 44GB of on-chip SRAM, keeping model weights resident on-chip and bypassing the memory-bandwidth bottleneck of conventional inference hardware — previously, speeds like this appeared only in small open-source models; this is a first for a frontier flagship. According to benchmarks officially released by the two companies: the full 2,500-question “Humanity’s Last Exam” took just 11 hours to complete, while the same exam took Claude Fable 5 more than three days; on the knowledge-work benchmark GDP-Val, end-to-end speedup is 5.6x.
SUPPLYIn agent scenarios where a single task strings together dozens of model calls, latency itself is the product — this tier is built for them. The bigger signal is in the supply chain: OpenAI has moved production inference for the frontier model onto non-Nvidia chips, and what Cerebras gets is a flagship-model endorsement, worth more than any benchmark score.
▪ SIGNALFor the first time, a frontier model runs on non-Nvidia chips as an official service layer — opening a gap in the procurement landscape for inference hardware.