OpenAI's Jalapeno Chip and What It Means for AI App Costs
OpenAI's new Jalapeno inference chip beat Nvidia's Blackwell on independent throughput-per-watt benchmarks. Here is what that could mean for AI app inference costs.
OpenAI has published its first independently verified benchmark numbers for Jalapeno, the custom inference chip it built with Broadcom, and the results hold up against Nvidia's current flagship. In tests run by the semiconductor research firm SemiAnalysis at the Hot Chips conference this week, Jalapeno delivered 1.5 to 1.9 times more AI inference work per watt than Nvidia's GB300 Blackwell system, with end to end latency cut by as much as 3.6 times on some workloads. For anyone building an AI powered product, the interesting part is not the chip itself. It is what happens to inference costs if numbers like these hold up once the chip actually ships.
What OpenAI actually showed
Jalapeno is OpenAI's first purpose built inference chip. It only runs trained models, it does not train them, and it was developed with Broadcom over roughly sixteen months before this week's public debut at Hot Chips in Palo Alto. SemiAnalysis, which operates its own independent hardware lab and published its own write-up of the results, tested Jalapeno against Nvidia's GB300 Blackwell system across three large models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, using the InferenceX benchmark suite built for cross-vendor comparisons like this one.
The numbers, and an important caveat
The core result: Jalapeno's 700 watt package produced 1.5x to 1.9x more throughput per watt than the GB300's 1,400 watt package at peak throughput, and OpenAI says performance on interactive, latency sensitive workloads came in 2.1x to 4.1x higher, with end to end latency 1.7x to 3.6x lower. TechCrunch's coverage of the benchmarks frames this as OpenAI's strongest public evidence yet that it can build competitive inference silicon in house.
One caveat matters. Jalapeno uses newer HBM4 memory, while the GB300 it was measured against still runs older HBM3e, which the-decoder's analysis and SemiAnalysis both flag as an unequal comparison. Nvidia's upcoming Vera Rubin platform also moves to HBM4, and SemiAnalysis has said that matchup, not yet run publicly, would be the fairer test. Jalapeno has also not adopted multi-token prediction or speculative decoding, techniques some competing systems already use to shave latency further. This is a genuine result, not a marketing slide, but it is a snapshot against last generation Nvidia hardware, not the generation it will actually compete with in production.
Why it matters if you build with AI
Nothing changes in API pricing today. OpenAI says Jalapeno is headed for low volume production only in late 2026, so it will not power meaningful ChatGPT or API traffic for some time. What the benchmark signals is more useful for anyone planning around AI costs than the number itself.
Inference, not training, is where the ongoing cost of an AI feature actually sits once a model ships, as covered in this breakdown of what inference costs and why. A chip that does more work per watt lowers the hardware and power cost behind every token, and that cost either gets passed to customers as savings or gets kept as margin.
Reasoning models, the kind that spend extra compute working through a problem before answering, covered in our explainer on reasoning models, are exactly the workload purpose built inference chips like this are aimed at. If Jalapeno-level gains reach production, reasoning heavy calls are the more likely place to see cheaper, faster pricing first.
This is one move in a wider contest. Nvidia is already fielding pressure from other directions, including Nvidia's own inference chip talks with Rebellions and the broader shift toward specialized NPU silicon for inference workloads. More credible competitors tend to move prices over the following product cycles, not overnight.
If your monthly spend on an AI coding agent or similar tool is mostly inference calls, as broken down in this look at what an AI coding agent actually costs per month, the honest takeaway today is that nothing changes yet. Watch for a real Rubin comparison and actual production deployment before expecting any of this to show up on an invoice.
What would make this matter more
Three things would turn this from a conference result into something that actually moves prices: an independent Jalapeno versus Vera Rubin benchmark on equal HBM4 footing, OpenAI shipping Jalapeno into real production inference capacity rather than engineering samples, and any sign that efficiency gains get passed through to API pricing rather than absorbed as margin. None of those have happened yet. Until they do, this is a credible signal that the inference hardware market has a real new competitor in it, worth tracking the same way you would any other fast moving AI story, not yet a reason to rebuild a cost model around.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


