First benchmarks put OpenAI’s Jalapeno ahead of Nvidia on efficiency

OpenAI's in-house inference chip beat Nvidia's GB300 on efficiency in tests run at Hot Chips.

ChipNews Staff
2 Min Read

OpenAI released the first public benchmark numbers for Jalapeno, its in-house AI accelerator built with Broadcom, and the results put the newcomer ahead of Nvidia’s current flagship on efficiency. The results were presented at Hot Chips and apply to inference workloads alone.

Jalapeno, a 700W part, delivered 1.5x to 1.9x more throughput per kilowatt than Nvidia’s GB200 and GB300 rack systems, which are rated at 1,200W and 1,400W. End-to-end latency came in 1.7x to 3.6x lower across three open models: GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI’s 1-trillion-parameter Kimi K2.5. OpenAI reported its widest leads at low-latency settings, where it claims up to 104x more throughput per kilowatt than a GB300 tuned for speed.

SemiAnalysis, which built the InferenceX benchmark suite and ran it alongside OpenAI engineers, said the chip beats every Nvidia, AMD and Google part it has tested. OpenAI plans to deploy Jalapeno in its own data centers later this year, and the design pairs a compute die with six HBM4 stacks totaling 216 GiB at 15.4 TB/s.

The results carry caveats. Jalapeno was not tested against Nvidia’s Vera Rubin platform, it cannot train models, and comparisons ran against GB300 systems using single-token prediction. Package TDP served as the normalization basis, while sustained draw measured below 550W.

Built in nine months from RTL to tapeout on a TSMC 3nm-class process, Jalapeno gives OpenAI leverage over the HBM supply Nvidia dominates. A second generation is approaching tapeout, and scaling the chip across OpenAI’s 10GW deal with Broadcom would make the company a major new claimant on HBM4.

Share This Article