AGI Soon As Possible · Deep reads on AI & tech
Article

OpenAI Unveils 'Jalapeño,' Its First Custom Inference Chip with Broadcom

2026-07-02 · 3 min read

On June 24, 2026, OpenAI unveiled Jalapeño, its first custom inference chip co-designed with Broadcom, and said large-scale deployment will begin at gigawatt scale in late 2026. The design-to-tape-out cycle took nine months, and OpenAI's own models helped accelerate parts of the design.

The Anatomy of the Announcement: Who Owns What

On June 24, 2026, OpenAI announced Jalapeño, its first custom chip co-designed with Broadcom. Jalapeño is a chip optimized for large language model (LLM) inference: OpenAI designed it, Broadcom handles manufacturing and networking, and Celestica handles boards, racks, and system integration. OpenAI described the chip as its first "Intelligence Processor." The division of labor is the tell here. OpenAI keeps the design reins while handing fabrication and assembly to specialists in each field, following the classic fabless playbook.

Why Aim at Inference, Not Training

OpenAI's choice to point its first chip at inference rather than training makes sense once you look at the shape of the cost. Training is a large, one-time expense when a model is built, but inference recurs with every user request for as long as the service runs. With users now numbering in the hundreds of millions, OpenAI's heaviest burden has arguably already tilted toward inference rather than training. Leading with performance per watt fits this context: power efficiency directly sets the size of that recurring bill.

How to Read the Numbers

OpenAI self-reported that Jalapeño's performance per watt is "substantially" better than current hardware. It did not release specific benchmark figures, however, and OpenAI promised a separate technical report. Large-scale deployment is set to begin at gigawatt scale in late 2026, and according to press reports, Microsoft is expected to purchase roughly 40% of the initial volume. The path from design to tape-out took nine months, and engineering samples are already running real workloads including the GPT-5.3-Codex-Spark model. The first-wafer handover included OpenAI's Sam Altman and Greg Brockman, along with Broadcom's Hock Tan and Charlie Kawwas.

Reorder those facts and the signal sharpens. Microsoft taking roughly 40% of the initial volume suggests the chip may spread beyond OpenAI's own consumption to become a standard part in partner data centers. The nine-month cycle, too, is aggressive against typical chip-design timelines.

What It Asks of Korea's Industry

The fact that performance per watt is the crux cuts both ways for Korea's semiconductor sector. On one side sits the worry that a wave of custom silicon could erode demand for general-purpose GPUs; on the other, the expectation that demand for high-bandwidth memory (HBM) and advanced packaging only grows, since inference chips still need memory and networking. For domestic fabless firms and foundries, the "service company designs, specialists manufacture" structure on display here is worth watching for how it may reshape future collaboration.

ASAP's View: What Remains Unverified

Jalapeño's performance edge is, for now, OpenAI's own claim. A "substantially better" phrasing with no benchmark figures sits somewhere between marketing and measurement, and it is safer to treat it that way. Gigawatt-scale deployment is likewise a forecast, not a result; whether the late-2026 date holds and whether early yields hit their targets will have to be confirmed by the separate technical report and real operating data. Until then, this announcement reads best as a statement of direction.


Source: OpenAI

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts