Skip to content

OpenAI previews GPT-5.6 Sol Ultrafast mode at 750 tokens per second via Cerebras wafer-scale chips

· by Pondero Newsdesk

The short version

OpenAI opened a limited API preview of Ultrafast mode on August 13, 2026, delivering up to 750 output tokens per second for GPT-5.6 Sol through Cerebras wafer-scale silicon, 14 times faster than standard processing with no reduction in model intelligence.

OpenAI previews GPT-5.6 Sol Ultrafast mode at 750 tokens per second via Cerebras wafer-scale chips

At 750 output tokens per second, GPT-5.6 Sol Ultrafast can answer a 2,500-question graduate-level benchmark in just over 11 hours. The same workload takes Claude Fable 5 more than three days. On August 13, 2026, OpenAI opened a limited API preview of the Ultrafast serving tier, powered by Cerebras (NASDAQ: CBRS) wafer-scale chips, and said the speed gain carries no reduction in the model's intelligence.

What

GPT-5.6 Sol Ultrafast launched in limited API preview on August 13, 2026, delivering up to 750 output tokens per second, per OpenAI's announcement. That is 14 times faster than the standard GPT-5.6 Sol serving tier. A small group of API customers received access at launch; businesses not yet in the pool can join a waitlist through Cerebras' website as capacity expands.

The hardware powering Ultrafast is Cerebras' Wafer-Scale Engine. Conventional GPU inference must repeatedly shuttle model weights between on-chip memory and off-chip storage. Cerebras' architecture keeps all weights resident in 44 GB of on-chip SRAM per wafer, per Cerebras' press release. That design eliminates the memory-bandwidth bottleneck that constrains inference speed on GPU clusters. For large frontier models with the most weight parameters to transfer, the gap is largest.

Per both the OpenAI announcement and the Cerebras press release, the intelligence of the model stays identical to standard GPT-5.6 Sol. No capability is traded away for the speed increase.

Target use cases named by OpenAI include live voice experiences, customer support automation, financial market analysis, e-commerce workflows, and security incident response. The company framed the launch around a directional claim: "Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second."

Why it matters

Operators building latency-sensitive products have long faced a forced tradeoff. Smaller, faster models ran at low latency but carried a quality ceiling. Large frontier models ran slowly enough to block the use cases that need sub-second responses. Ultrafast directly challenges that framing. If the intelligence parity claim holds in third-party evaluation, the design space for voice agents, real-time financial systems, and automated triage expands.

Live voice assistants require a response within roughly 200 milliseconds to feel natural. Customer support pipelines handling peak load cannot afford to queue. Security tooling running triage during an active incident needs to keep pace with incoming signal volume. Each of those applications has historically forced engineering teams to a smaller model and the quality ceiling that came with it.

750 tokens per second is also fast enough to shift the economics of batch workloads. Running Humanity's Last Exam in 11 hours instead of three days is not just a benchmark curiosity. It maps to document review, contract analysis, regulatory examination, and any enterprise use case where a fixed set of questions needs to be answered against a large corpus. Throughput at this level turns those from overnight jobs into same-session tasks.

Pricing has not been disclosed. Cerebras inference has historically carried a premium over GPU-based serving, and whether OpenAI absorbs that cost into existing API tiers or introduces a separate Ultrafast price point is unknown. That figure will shape adoption outside the preview cohort more than any benchmark result.

Anthropic offers Claude Fast mode as the nearest existing alternative, but it does not deliver comparable output speeds, per TechCrunch's coverage.

Context and reactions

The benchmark numbers in Cerebras' press release give the most concrete picture of what 750 tokens per second changes in practice. On Humanity's Last Exam, a 2,500-question set spanning graduate-level chemistry, economics, and literature, GPT-5.6 Sol Ultrafast completed the full set in just over 11 hours with accuracy comparable to Claude Fable 5. Claude Fable 5 required more than three continuous days, placing Ultrafast roughly 7x faster on that workload, per Cerebras' announcement. On GDP-Val, a benchmark of knowledge-work tasks covering legal briefs, financial models, and engineering reports, Ultrafast delivered a 5.6x end-to-end speedup with no quality loss.

Cerebras also cited Artificial Analysis data for head-to-head speed comparisons: Ultrafast runs 5x faster than Claude Opus 4.8 in Fast mode and 11x faster than Claude Fable 5. Those are vendor-attributed figures citing a third-party aggregator, not independent lab benchmarks. Developers in the preview cohort will be the first to confirm or revise those comparisons in real workloads.

Andrew Feldman, CEO and co-founder of Cerebras, said in the press release: "GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive. Together with OpenAI, we're putting frontier intelligence in the hands of users at unprecedented speed and changing what's possible with AI."

Sachin Katti, VP Compute Strategy and GPT-Infra at OpenAI, described the launch as a deliberate learning phase: "By combining GPT-5.6 Sol with Cerebras' inference technology, we're exploring what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency. We're starting with a small group of customers to learn where that speed creates meaningful value, and we'll use those learnings to inform how we expand the service over time."

Cerebras went public on NASDAQ under the ticker CBRS. Based in Sunnyvale, California, the company's Wafer-Scale Engine architecture predates the current large-model inference market by several years. The key design choice is that a single wafer-sized chip avoids the inter-chip communication overhead that slows inference on GPU clusters. When a frontier model's weights fit entirely in on-chip SRAM, every generated token avoids repeated round-trips to off-chip HBM memory. The efficiency gains compound with model scale, which is why the architecture is especially suited to GPT-5.6 Sol's weight footprint.

What to watch next

Pricing disclosure is the most actionable near-term signal. OpenAI has not set public Ultrafast rates, and Katti's statement suggests the commercial structure is still forming around early customer feedback. If Ultrafast lands near standard GPT-5.6 Sol API pricing, Anthropic and Google will face direct pressure to announce comparable high-speed inference tiers. A significant premium would narrow adoption to latency-critical enterprise workloads where the speed justifies the cost.

The second signal worth tracking is whether Cerebras signs additional frontier lab partnerships. A wafer-scale arrangement with Anthropic or Google DeepMind would move this hardware approach from an OpenAI-exclusive to a cross-industry infrastructure standard, with real consequences for how the inference market prices out over the next 12 months.

Sources