OpenAI’s Jalapeño beats Nvidia Blackwell on speed and efficiency
OpenAI’s custom inference chip, codenamed Jalapeño, showed benchmark results that beat Nvidia Blackwell systems on latency and energy efficiency across several open models. The chip is not yet deployed at Nvidia scale, but this is one of the clearest signs that the biggest AI labs are turning compute cost into a silicon-design problem.
Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.https://t.co/aej6vABD9q
— TechCrunch (@TechCrunch) August 25, 2026
Q1What was actually announced?
The benchmark result was presented at Hot Chips and reported by TechCrunch in its summary. Jalapeño was tested on SemiAnalysis InferenceX across models including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5, with comparisons against Nvidia GB200 and GB300 systems.
Q2How big is the signal?
Reported results show roughly 1.5 to 1.9 times more AI work per watt and about 1.7 to 3.6 times lower end-to-end latency in selected tests. OpenAI developed the accelerator with Broadcom. Limited deployment is expected around the end of 2026, with larger-scale use planned in 2027.
Q3Why does it matter now?
For a frontier lab, inference is one of the largest recurring expenses. Even a modest efficiency gain can save billions of dollars at massive token volume. Custom chips also give OpenAI more negotiating leverage with Nvidia and let hardware be tuned around its own models, networking, and serving software.
Q4What is the catch?
Benchmark wins do not equal a complete platform win. Nvidia bundles GPUs with CUDA, networking, mature libraries, supply, and a huge developer ecosystem. Jalapeño is an internal accelerator with limited deployment, and the published comparisons reflect chosen workloads. Scaling manufacturing and software may be harder than designing a fast chip.
Q5What should we watch next?
Watch production volume, what share of OpenAI inference moves onto Jalapeño, Broadcom capacity, and whether later chips train models as well as serve them. Nvidia faces real pressure when custom silicon takes a meaningful percentage of frontier-lab compute spending, not simply when it wins a benchmark.
