OpenAI publishes Jalapeño's first measurements: 1.5x to 1.9x more work per watt, up to 3.6x lower latency
OpenAI published the first measured results for Jalapeño, its custom inference chip, on August 25, 2026, reporting 1.5x to 1.9x more AI work per watt and 1.7x to 3.6x lower end-to-end latency than the comparison systems across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. For highly interactive workloads the chip delivered 2.1x to 4.1x higher performance. Measurements ran on InferenceX, a public benchmark from SemiAnalysis, against NVIDIA GB200 and GB300 systems. OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of the year, with a second generation deep in development and a third taking shape. ASAP summarizes the result from OpenAI's announcement and its published appendix as primary sources.
The number that was missing in June has arrived
OpenAI unveiled Jalapeño with Broadcom on June 24, 2026, claiming only that performance per watt was "substantially" better than existing hardware and publishing no benchmarks at all. That announcement promised a separate technical report, and the document released on August 25 fills the gap. The difference in character between the two is clear: June was a statement of direction, August is a set of measurements.
Something else changed. The June announcement spent its length on the chip, its partners and the schedule, while this one leads with the conditions under which the numbers were taken. OpenAI defines its evaluation as performance "at a matched user experience," meaning how much useful AI work each system completes per unit of power while meeting the latency that customers and interactive agents require. The company adds that agents run many steps in sequence, so delay at each step compounds across an entire task.
What the three public models actually produced
Jalapeño recorded 85,448 mixed tokens per kilowatt on GPT-OSS 120B against 44,960 for the GB200 comparison system, roughly 1.9x. End-to-end latency in the same run was 1.03 seconds versus 1.80, about 1.7x lower, and minimum time-between-tokens was 0.69 milliseconds versus 1.87, about 2.7x shorter. Converted to per-user throughput that is 1,459 tokens per second against 535.
The gap widened on DeepSeek R1 670B run in MXFP4. Peak mixed throughput came in at 19,641 tokens per kilowatt against 11,781, roughly 1.7x; end-to-end latency at 1.65 seconds against 5.99, roughly 3.6x; minimum TBT at 1.43 milliseconds against 5.90, roughly 4.1x. The comparison system there is the GB300. On Kimi K2.5 1T, the largest public model tested, Jalapeño reached 18,195 tokens per kilowatt against 11,862, roughly 1.5x, with latency of 1.56 seconds against 5.31, roughly 3.4x, and minimum TBT of 1.44 milliseconds against 5.48, roughly 3.8x. Every measurement used a nominal 8k input and 1k output.
Lining the three models up reveals a pattern. As the model grows, the perf-per-watt multiple shrinks from 1.9x to 1.5x while the latency multiple expands from 1.7x to 3.4x. Throughput advantage dilutes with scale; latency advantage compounds with it. OpenAI's added note that the advantage widened further on its own frontier models points the same direction.
The arithmetic that puts 700 watts next to 1,400 watts
OpenAI's appendix states that results were normalized using published package TDP: 700 watts for Jalapeño, 1,200 watts for the GB200, 1,400 watts for the GB300. That single line is the most important caveat in the release, because performance per watt is a metric whose multiples move entirely with the choice of denominator.
The same document reports that Jalapeño's measured sustained power stayed at or below 550 watts on the workloads tested, more than 21% under its 700-watt rating, yet 700 watts is what the normalization used. That choice cuts against Jalapeño and is conservative on its own terms. The problem is that measured power for the comparison systems was not published. Without knowing how much the GB200 and GB300 actually drew against their 1,200-watt and 1,400-watt ratings, the exact size of the multiple stays unsettled. If those systems also ran below their ratings, the gap narrows from what is reported here.
None of this makes rating-based normalization an illegitimate method. Data center power is provisioned against rated draw rather than measured draw, so for the question of how many accelerators fit in a rack, the rating is the realistic denominator. The meaning shifts, though, the moment the number is restated as "how much work you get from the same electricity." Those are answers to two different questions, and this release stands on the first.
Where the 53.7x and 104.3x figures come from
The largest multiple in the appendix, 104.3x, comes from the DeepSeek R1 comparison with the operating point pinned at 169.41 tokens per second per user. That point is the fastest response rate the comparison system can produce, the GB300's own minimum TBT. The same construction yields 53.7x on GPT-OSS 120B and 56.1x on Kimi K2.5.
The anatomy of those double-digit multiples is straightforward. When the comparison system runs at its top speed, almost no throughput is left over: 118 tokens per kilowatt in the GB300 case. Jalapeño sustains the same speed while processing 12,258 tokens per kilowatt. The comparison captures a point at the far edge of the curve where one side falls off sharply, so restating it as an overall chip-to-chip difference overstates the result.
That does not make the figure meaningless. Interactive agents live exactly in that region. A response a user waits on, or an agent loop firing tools back to back, demands response speed before throughput, and how many concurrent users you can hold at that speed is what sets the cost of the service. The real question the table poses is not the size of the multiple but whether throughput survives in the high-speed region. OpenAI's own summary makes the same point: ultra-fast-mode inference at efficiencies previously available only in fast mode, and fast-mode inference at efficiencies previously available only in batched mode.
Porting kernels in two months is the bigger story
OpenAI used Codex with GPT-Astra to bring three open-weight models that were never part of Jalapeño's original production plan to high performance within two months. On selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5x to 1.8x faster than the existing human-expert-written versions. OpenAI states explicitly that those figures apply to the selected blocks and not to the full model.
Even with that qualifier attached, this paragraph outweighs any perf-per-watt multiple. New accelerators usually lose in the market on software rather than silicon. Every new model family demands fresh kernels and model-specific optimization, and while that work consumes quarters of human engineering time the incumbent ecosystem keeps its lead. This release presents a case where that cycle compressed to two months. When switching costs fall, the threshold for choosing different hardware falls with them.
Jalapeño's design points the same way. OpenAI describes the chip as a "clear, predictable programming target for both humans and AI," where engineers express work through local tensors, explicit communication and predictable synchronization, and AI then optimizes how that work is mapped, placed, scheduled and coordinated across the system. The goal is to reshape parallel programming into a problem AI can tackle. The nine months from design to tapeout came from the same source, with AI shortening design, measurement and verification loops and optimizing the chip's arithmetic circuits.
What this means for memory suppliers and power-constrained data centers
OpenAI's argument that performance per watt is the better standard mirrors the 2026 reality of data centers, where power rather than silicon is the binding constraint. Stating that per-unit-of-power performance is more useful than per-chip performance is a practical judgment made in an environment where sites and substation capacity run out first. Korean operators face the identical constraint, with permitting and grid connection setting the pace of the business.
The release carries a split signal for memory suppliers. Jalapeño's core design idea is to place model state explicitly, including the KV cache used while generating a response, keep it local and minimize data movement, treating prefill as compute-bound and decode as memory-bandwidth-bound and building for both phases in one architecture. An architecture whose performance turns on memory bandwidth and locality does not reduce the importance of high-bandwidth memory or advanced packaging. Custom silicon may erode demand for general-purpose accelerators while the memory curve moves separately.
The release also restates the entry requirement for custom silicon. What OpenAI holds is not design capability alone but a position from which models, products, serving software, chips, memory, networking and systems can be designed together, with real workloads feeding improvements back into every layer. That structure is hard to copy for anyone strong in only one layer. The transferable lesson is the span of the vertical integration, not the chip specification.
What remains unverified
Jalapeño's numbers were measured on InferenceX, a public benchmark from SemiAnalysis, but the measurement and the reporting were both done by OpenAI. Using a public benchmark is a step beyond a purely in-house metric, and it is still not third-party verification. The appendix conditions are a nominal 8k/1k profile at STP, and whether the same multiples hold at other input and output lengths or under different concurrency is outside the scope of this document.
Deployment remains a forecast. OpenAI says it will begin deploying Jalapeño inside its own infrastructure by the end of the year while noting that production qualification, software maturation, preparation to operate at scale and validation across more models are all still under way. No sales plan or pricing is mentioned. Jalapeño is not a product competing with NVIDIA in the market but a component that changes OpenAI's internal cost structure, and the same document makes that plain by committing to continue deploying accelerators from NVIDIA and other partners widely for both training and inference.
One question goes unanswered. How much of a 1.5x to 1.9x perf-per-watt gain converts into a percentage reduction in the cost of serving depends on chip price, yield, rack density and operational efficiency together. OpenAI writes that it expects improved operating leverage, with useful work and revenue growing faster than the cost to serve, but the number that would substantiate the size of that leverage is not in this document. The next checkpoint is real operating data after deployment begins at the end of the year.
Source: OpenAI, "Jalapeño's first results show industry-leading speed and efficiency in AI inference" (August 25, 2026), including its published appendix. Summarized by ASAP.

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr