❯ OpenAI Reveals First Benchmarks for In-House Inference Chip Jalapeño — Per-Watt Performance Beats Nvidia GB300
FIRST BENCHAt the Hot Chips conference on August 25, OpenAI released the first performance data for its in-house inference chip Jalapeño: across three models — GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T — effective compute per watt reaches 1.5x to 1.9x that of Nvidia’s GB200 and GB300 systems, with end-to-end latency 1.7x to 3.6x lower. The chip was co-designed by OpenAI and Broadcom, with tests running on SemiAnalysis’ InferenceX platform.
HARDWARE CARDThe single package carries 6 HBM4 stacks — 216 GiB of capacity and 15.4 TB/s of bandwidth — rated at 700 watts, yet under measured loads it stays under 550 watts. The rival GB300’s full-card draw is 1,400 watts. Under highly interactive workloads, the gap widens further to 2.1x to 4.1x. A first-gen chip beating Nvidia’s current flagship on this curve is something the past three years haven’t seen.
CAPACITY KEYOpenAI’s own deployment cadence is restrained: only a small-batch ramp by the end of 2026, with full scale-out waiting until 2027. So in the near term, Nvidia won’t lose a single unit of shipments — the real change is at the procurement table. When a model company holds an inference chip it can build itself and that’s cheaper per watt, the bargaining chip changes hands. In power-constrained data centers, tokens per watt convert directly into revenue — the only metric OpenAI cares about.
▪ SIGNALA first-gen ASIC beating a current flagship GPU isn’t a process-node win — it’s the ability to put models, compilers, and silicon into the same iteration cycle.