Companies
OpenAI's Jalapeño Chip Outruns Nvidia in Inference Tests
OpenAI's first custom chip, built with Broadcom, beat Nvidia's GB200 and GB300 on performance per watt and latency in partially verified inference benchmarks.

OpenAI now has working silicon that beats Nvidia at inference, and for once the claim does not rest on the vendor's word alone. On August 25 the company published first results for Jalapeño, the custom chip it has been building with Broadcom since mid-2024. Measured with SemiAnalysis's InferenceX benchmark suite, with part of the testing verified on site by the analyst firm, Jalapeño delivered 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 systems.
The rest of the numbers follow the same pattern. End-to-end latency came in 1.7 to 3.6 times lower, and interactive chat-style workloads ran 2.1 to 4.1 times faster. The tests covered open-weight models of very different sizes: GPT-OSS 120B, DeepSeek R1, and the trillion-parameter Kimi K2.5. On GPT-OSS 120B, the chip pushed roughly 1,400 tokens per second.
Nine months from design to fab
The timeline is almost as striking as the benchmarks. Design work started in mid-2024 and the chip taped out for production in November 2025. SemiAnalysis founder Dylan Patel, not a man known for handing out compliments, noted that first-generation chips are usually uncompetitive, yet this one beats Blackwell and even the newer Rubin generation. Two caveats belong next to that praise. Jalapeño is still at the engineering-sample stage, with broad deployment expected in 2027, while Nvidia is already shipping Vera Rubin systems to customers. And OpenAI chose which models and scenarios were tested; the independent verification was partial.
A specialized chip beating a general-purpose GPU is less surprising than it sounds. Nvidia's hardware has to stay ready for every workload, training included. Silicon dedicated to one job, running a finished model, can organize its memory access and data flow around that single pattern, trading flexibility for efficiency. The timing was hardly accidental either: the announcement landed during Hot Chips week, the conference that sets the industry's silicon calendar, the same week Nvidia detailed its new Vera CPU architecture.
What it means for Nvidia
Not much in the short term. OpenAI remains one of Nvidia's largest customers, and Nvidia recently agreed to backstop OpenAI's Ohio data center lease with up to $105 billion. Jalapeño also targets inference only, not training. The longer-term signal is harder to dismiss: every major lab wants its most expensive workload on its own silicon. Anthropic is building a chip design team, AMD bought Taalas to bake models directly into hardware, and Waymo designed its own robotaxi chip. Custom inference silicon is becoming table stakes at the frontier.
The part that reaches your invoice
Inference cost is the bill every AI-using business ultimately pays; API pricing sits on top of it. When chips like this raise efficiency per watt, providers gain room to cut prices, and they have been using that room: OpenAI trimmed GPT-5.6 Sol's API price by 20 percent earlier this month. If you are budgeting a chatbot, a document pipeline, or an agent rollout, the unit economics you calculate today will likely look conservative in two years. One thing efficiency gains will probably not deliver is lower total energy use. Cheaper inference has a long track record of producing one outcome above all: more inference.
Sources: OpenAI, The Decoder, TechCrunch

Written by
Faruk Talmaç
Co-Founder & Editor
Co-founder of YZ Uzman, with 20+ years of experience in web design and software development.