The Kitchen Watt
Live
Cocina

OpenAI chip beats Nvidia in efficiency

OpenAI has revealed performance benchmarks for its custom Jalapeño AI chip, claiming it delivers up to 1.9x better throughput per kilowatt than Nvidia's

OpenAI has revealed performance benchmarks for its custom Jalapeño AI chip, claiming it delivers up to 1.9x better...

OpenAI published the first performance figures for its custom Jalapeño AI chip on August 25. The company claims the chip, co-developed with Broadcom, delivers significantly better efficiency than comparable Nvidia hardware for running AI models.

According to OpenAI, the Jalapeño inference accelerator provides 1.5 to 1.9 times more throughput per kilowatt than Nvidia's GB200 and GB300 rack systems. It also claims end-to-end latency is 1.7 to 3.6 times lower. These figures were presented at the Hot Chips conference.

Reported Performance Benchmarks

OpenAI's benchmarks compared the Jalapeño chip against Nvidia's GB200 and GB300 systems using three open-source AI models. The tests used a nominal 8,000-token input and 1,000-token output. The company's reported efficiency gains vary by model.

Model TestedCompeting Nvidia SystemClaimed Efficiency Gain (Throughput per kW)
GPT-OSS 120BGB2001.9x
DeepSeek R1 670BGB3001.7x
Moonshot AI's Kimi K2.5GB3001.5x

Richard Ho, who leads OpenAI's hardware program, called the results "a very, very significant performance advance over state-of-the-art." The Jalapeño chip is rated for a 700W power draw. The Nvidia systems it was compared against were rated for 1200W and 1400W.

Technical Specifications and Context

The Jalapeño is an Application-Specific Integrated Circuit (ASIC) designed primarily for AI inference workloads. OpenAI says it built the chip to keep model state local, minimizing data movement and communication delays.

Each chip package pairs a compute die with six HBM4 memory stacks, providing 216 GiB of capacity at 15.4 TB/s of bandwidth. This is less capacity than the GB300's 288GB of HBM3E memory but offers more bandwidth. It also provides roughly 50 percent more memory per watt of rated power.

All published performance numbers are from the A0 silicon, which is the initial version of the chip. A more efficient B0 stepping is already in fabrication, reportedly offering about 25 percent better performance per watt.

Limitations and Competitive Landscape

Analysts note the announcement's significance. Alexander Harrowell of Omdia told CNBC that "this is the biggest competitive threat to NVIDIA." He pointed out that roughly half of AI infrastructure spending comes from firms that already run or could run custom chip programs. Adrien Sanchez of Yole Group framed it as pressure on Nvidia's inference margins specifically.

However, the benchmarks have important caveats. Nvidia's GB300 still leads on absolute throughput per package by 20 to 25 percent. The efficiency comparison is made on a per-kilowatt basis, not a per-chip basis. When OpenAI pits Jalapeño against a GB300 configured for multi-token generation-common in Nvidia's production deployments-the peak efficiency lead drops to roughly 1.5x.

OpenAI has not published any information comparing Jalapeño to Nvidia's upcoming Vera Rubin architecture, which Nvidia has presented as an efficiency leader. The Jalapeño chip is also limited as an ASIC useful mainly for inference, while Nvidia's chips remain key for training new AI models.

Volume production of the Jalapeño chip is not expected to ramp up until 2027. The timing of the announcement is notable, coming as Nvidia just agreed to backstop $105 billion in financing for OpenAI's upcoming Ohio facility. Richard Ho told Bloomberg that OpenAI continues to need a lot of Nvidia's chips, calling the company "a really good partner."

Topics

#Cocina

Related coverage

More from Cocina