2026-08-19-Wed · Cerebras

From Issue 18 (2026-08-19) · 14 stories in this issue

❯ Cerebras launches CS-4 rack with three wafer-scale chips, first deliveries begin this quarter

PRODUCT LAUNCHCerebras has released the rack-scale system CS-4, packing three WSE-3 Turbo wafer-scale chips into a single unit, with first deliveries beginning this quarter. The company says per-user token output can reach up to 30 times that of GPU-based solutions, and calls it the industry’s fastest AI accelerator. Each WSE-3 Turbo still packs 4 trillion transistors, 900,000 AI cores, 46,225 square millimeters of silicon, and 44GB of on-chip static memory.

ARCHITECTURECS-4 is the first product on the new Nexus platform architecture, modular across three domains: compute, power delivery, and I/O. The most revealing choice is power: power conversion now sits roughly 0.5 mm from the processor, versus about 50 mm on conventional GPU boards — effectively erasing board-level losses. The programmable I/O subsystem supports two connection modes, doubles bandwidth, and cuts latency from 5 microseconds on the prior generation to as low as 2 microseconds.

SPECSPer company disclosures, single-chip compute rises to 250 PFLOPS, with memory bandwidth of 43.2 PB/s, on-chip interconnect of 53.5 PB/s, and off-chip I/O of 2.4 Tb/s — all double the previous WSE-3 generation. In Cerebras’ earlier published comparisons, the CS-3 ran gpt-oss-120B at more than 2,700 tokens per second, versus about 900 for Nvidia’s B200 on the same task.

STRATEGYCerebras’ bet has never been training; it’s per-user token throughput — the single metric that decides whether users in chat and agent scenarios feel there’s “no waiting.” Teams building real-time agents now have to recalculate the exchange rate between latency and unit price: the wafer-scale approach carries a higher unit price and a narrower ecosystem, but in long-chain tasks, the wait saved at each step gets amplified by step count. Whether that arithmetic works in their favor won’t be answered until the first units reach customer data centers.

▪ SIGNALMoving power delivery to 0.5 mm from the chip is itself the message: this generation’s bottleneck is no longer compute density — it’s how to get the power in.