OpenAI published first Jalapeño results yesterday, and the numbers are spicy enough to dominate the news cycle. The company claims 1.5 to 1.9 times more AI work per watt than Nvidia’s GB200 and GB300 systems. Up to 3.6 times lower end-to-end latency. Hundreds of tokens per second for a single user running DeepSeek R1.
The headlines will call it a knockout. It isn’t. What we are looking at is a carefully controlled proof of concept that shows exactly where Nvidia is vulnerable, and just how fast AI-assisted chip design is compressing the silicon cycle.
The Benchmarks Are Real, But They’re Narrow
The figures carry more weight than a typical press release because SemiAnalysis engineers verified some runs on-site in OpenAI’s lab. That matters. But the test conditions are heavily optimized for Jalapeño’s specific strengths. The chip is inference-only. It uses HBM4 memory at 15.4 TB/s. The workloads were decoding-heavy, short-context runs across GPT-OSS 120B and Kimi K2.5 1T. And crucially, Jalapeño hits these numbers without multi-token prediction, speculative decoding, or prefill-decode disaggregation.
That last point is telling. Nvidia’s moat has never been raw throughput on a single idealized model. It is flexibility. Jalapeño is a custom ASIC co-developed with Broadcom that knows exactly what OpenAI’s serving stack looks like. When you control the model weights, the scheduler, the memory placement, and the silicon itself, you can push the latency-throughput frontier in ways a general-purpose GPU simply cannot. But that same specialization means the chip is useless for training. OpenAI still needs Nvidia or another vendor for that, and the company has been explicit that Jalapeño will not replace its broader chip strategy.
I spent time yesterday watching the technical community dissect the fine print, and the memory disparity was the immediate flashpoint. HBM4 versus HBM3E is not a minor footnote. It is a generational bandwidth advantage that explains a healthy chunk of the efficiency gap before you even get to architecture.
The chip was also tested without speculative decoding or multi-token prediction, which are standard tricks on Nvidia hardware. Jalapeño doesn’t need those shortcuts because the full stack is co-designed. That is a genuine architectural win, but it is also a luxury no general-purpose vendor can replicate.
The skepticism I saw focused on the gap between lab benchmarks and production-scale economics. These chips were not tested on agentic suites like AgentX, long-context memory pressure, or multimodal tool-use patterns. The impressive low-concurrency results, while fantastic for user-facing interactivity, don’t tell us what happens when thousands of these units hit a real fleet with noisy neighbors and variable context lengths. Everyone wants 700 tokens per second in a demo. Nobody wants 40 in production.
The Loop Is Closing
The part that should actually worry Nvidia isn’t the wattage chart. It is the calendar. OpenAI moved from initial RTL design to tape-out in nine months. That is not just aggressive. It is potentially a new paradigm.
The company used its own models to assist architecture exploration and optimization, which means AI is now designing the silicon it will eventually run on. I saw this framed online as the loop closing, and the description fits. Traditional ASIC cycles run 18 to 30 months. If OpenAI can iterate custom inference silicon on compressed timelines, it does not need to match Nvidia’s generality. It only needs to stay one generation ahead on its own workload. That changes the math for every hyperscaler watching from the sidelines.

This creates a strategic de-risking play that runs deeper than benchmark hero numbers. Whoever controls the chip controls the cost. With a $500B data center deal still hanging over the market and an Ohio $100B guarantee in the ground, OpenAI is sending a clear signal. It will buy Nvidia when it must. But it would rather not need to.
There are hard constraints here. Jalapeño is rated at 700W TDP but won’t reach volume deployment until 2027. Early samples running at or below 550W sustained are promising, but they are not a fleet.
There is no third-party sales path announced. Broadcom and TSMC still own the supply chain. And without speculative decoding or MTP, the efficiency gains rely entirely on brute memory bandwidth and tight stack integration. That works beautifully for ChatGPT-style chat. Whether it holds for long-horizon reasoning or multimodal agents is still an open question.
What Jalapeño proves is not that Nvidia is finished. It proves that the moat around general-purpose AI silicon is getting narrower every quarter. The real competitive edge is shifting from who fabricates the best transistor to who can close the feedback loop between model behavior and hardware geometry fastest. OpenAI’s first custom chip is already competitive on the metrics it cares about most: interactivity and inference economics at low concurrency. By the time this reaches real scale in 2027, the next generation will likely be in tape-out. That is the timeline that matters. For everyone else betting on a permanent Nvidia dependency, yesterday’s benchmarks were a small, sharp reminder that vertical integration is back in style.






