Skip to content

OpenAI's Jalapeno inference chip posts 1.5-1.9x more throughput per watt than Nvidia GB300 in first public results

· by Pondero Newsdesk

The short version

OpenAI published InferenceX benchmark results for Jalapeno on August 25, showing 1.5 to 1.9 times more throughput per kilowatt than Nvidia GB200 and GB300 systems at 550W sustained versus 1,200 to 1,400W for the Nvidia hardware. Small-volume deployment is planned for end of 2026.

OpenAI's Jalapeno inference chip posts 1.5-1.9x more throughput per watt than Nvidia GB300 in first public results

At 550 watts sustained, OpenAI's custom Jalapeno chip delivered 1.5 to 1.9 times more AI work per kilowatt than Nvidia GB200 and GB300 systems rated at 1,200 to 1,400 watts in the InferenceX benchmark, per TrendForce. The company published the results on August 25.

What the benchmark showed

SemiAnalysis ran the InferenceX tests across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Jalapeno cut end-to-end latency by 1.7 to 3.6 times versus the Nvidia systems, and 2.1 to 4.1 times for ultra-low-latency inference workloads, per The Register.

Each Jalapeno accelerator carries 216 GB of HBM4 memory at 15.4 TB/s bandwidth. A rack of 128 chips reaches 1.7 exaFLOPS of 4-bit compute and holds 27.5 TB of total HBM4 memory, per The Register. TSMC fabricated the chip on its 3nm process. Samsung Electronics reportedly supplies the HBM4 stacks, with Broadcom serving as the ASIC design partner, per TrendForce.

Jalapeno handles inference only. OpenAI still relies on Nvidia hardware for training runs.

Why it matters

Inference cost per token tracks directly to compute per watt. A chip that delivers 1.5 to 1.9 times more output per watt against the dominant inference accelerator class changes the cost structure for running ChatGPT and the GPT API at scale.

OpenAI plans small-volume deployment by end of 2026 and a broader rollout through 2027, per TechCrunch. GPT API pricing will not reflect these savings before 2027 capacity arrives, but the efficiency gap would give OpenAI a cost structure difficult for GPU-dependent competitors to match at that scale.

One caveat: OpenAI commissioned the InferenceX tests through SemiAnalysis. Third-party replication has not been completed. The benchmark also compared Jalapeno against current Nvidia hardware rather than next-generation successors to the GB300 that are in development.

Context

Jalapeno's development runs under OpenAI's 10GW chip deployment deal with Broadcom, signed in October 2025, per TrendForce. A second-generation chip is expected to reach tape-out within months, with third-generation development already underway per TrendForce. OpenAI's hardware lead Richard Ho described the chip as designed to minimize data movement and communication delays in the prefill and communication phases, per TechCrunch.

What to watch next

Third-party benchmark replication is the first credibility test for the performance numbers. Nvidia's response, whether in the form of updated inference chip specifications or timing announcements ahead of Jalapeno's 2027 volume ramp, is the competitive signal to track. If OpenAI adjusts API pricing after small-volume deployment begins at end of 2026, that is the earliest market signal the efficiency gains are reaching the cost structure.

Sources