Cerebras CS-4 stacks three wafers to reach 750 PFLOPs and 30x faster inference than GPU alternatives
At 4,400 tokens per second per user on a frontier 120-billion-parameter model, the CS-4 is not incrementally faster than GPU-based inference. It is fast enough to allow an agentic system to run more than an order of magnitude more reasoning steps in the same wall-clock time that a GPU-based system uses for a single pass. That is the concrete capability Cerebras put on the market on August 18 with the CS-4, a rack-scale accelerator built from three wafers under a new modular architecture called Nexus.
What happened
Cerebras unveiled the CS-4 on August 18, 2026, combining three WSE-3 Turbo wafers in a single rack system, per the company's investor press release. Each WSE-3 Turbo wafer contains four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon, with 44 GB of on-wafer SRAM. The three-wafer rack delivers 750 PFLOPs of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O bandwidth.
The 30x-over-GPU figure comes from Cerebras' own benchmark on GPT-OSS-120B, where the CS-4 reached 4,400 tokens per second per user. Cerebras also claims 10x more throughput per watt and 2x the speed compared with the CS-3.
Three architectural changes enable those results. First, Cerebras moved power conversion from roughly 50 millimeters away from the processors to 0.5 millimeters, effectively doubling power delivery to the wafers. Second, a pluggable backpack design physically decouples compute from the power array, cutting component count by 50 percent and reducing on-site deployment time from days to hours, per the same press release. Third, a programmable I/O module supports both standard RoCE v2 RDMA and a new Direct Wafer Links mode that connects wafers at two-microsecond latency without a network switch, enabling clusters that can serve models with more than 50 trillion parameters.
First CS-4 shipments are scheduled for Q3 2026. A footnote in the Cerebras press release states that actual throughput varies by model architecture, context length, precision, and serving configuration.
Why it matters
Sean Lie, co-founder and CTO of Cerebras, described the practical implication in the press release: a 30x speed advantage gives an agentic system room for more than an order of magnitude as much reasoning, verification, or tool use in the same wall-clock time as a GPU-based generation. For operators running multi-step agentic pipelines today, where a single turn with several tool calls can take seconds, that constraint largely disappears at 4,400 tokens per second per user.
The 10x throughput-per-watt gain changes data-center economics independently of the speed story. More tokens per watt means more revenue per rack. Dylan Patel of SemiAnalysis, quoted in the Cerebras announcement, called the CS-4's Backpack design "focused on time to market" and cited its ability to scale token production across larger models at volume.
The Direct Wafer Links capability is the less-discussed detail that carries the most forward-looking weight. Two-microsecond wafer-to-wafer latency without a switch opens the platform to 50-trillion-parameter-plus models, a range well beyond what current GPU inference clusters serve cost-effectively. Cerebras named AMD Helios and AWS Trainium as ecosystem partners for heterogeneous prefill/decode configurations that could take advantage of that bandwidth.
What to watch next
Independent benchmark data from CS-4 buyers will be the real test of the 30x claim. Shipments start this quarter, so third-party results should surface in Q4 2026 or early 2027. Heterogeneous disaggregated inference setups, pairing CS-4 with AMD Helios or AWS Trainium for the prefill stage, are the first integration patterns to follow once those ecosystem partnerships produce reference architectures.
Sources
- Cerebras Unveils CS-4: Up to 30 Times Faster than GPU-based Solutions: Cerebras investor press release, August 18, 2026
- Cerebras Overclocks WSE-3 Waferscale Engine To Boost Inference Oomph In Nexus CS-4: The Next Platform, August 19, 2026
- Cerebras Intros Faster WSE-3 Turbo Processor and First Rack-Scale CS-4 System: ServeTheHome, August 2026
