Almost every ChatGPT response delivered to millions of users runs on Nvidia custom silicon — at serious power cost and serious expense to OpenAI. That bill has been accumulating for years. Now OpenAI has built its own chip: Jalapeño, developed with Broadcom, designed specifically for inference (serving model responses to users, not building the model). The early benchmark numbers are striking. The catch: OpenAI ran the tests themselves.
What the Numbers Actually Say
The reported gains are large enough that infrastructure buyers should pay close attention — with one important asterisk.
Across SemiAnalysis’ public InferenceX benchmark, testing three models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — OpenAI reported 1.5x to 1.9x more AI work per watt at peak throughput versus comparison systems, and 1.7x to 3.6x lower end-to-end latency. For interactive workloads, where a user is actively waiting for a response, performance reportedly jumped 2.1x to 4.1x higher. These are not rounding errors.
A few supporting details worth holding:
- Chip rated at 700W; sustained measured power stayed at or below 550W during tested workloads
- Models tested span 120B to 1 trillion parameters — a serious range
- Broadcom CEO Hock Tan positioned it as comparable to Nvidia Blackwell and Google TPUs, according to Reuters
- OpenAI used Codex and GPT-Astra to shorten design and verification cycles — AI helping build AI hardware
The architecture insight matters here. Inference runs in two phases: prefill (compute-heavy) and decode (memory-bandwidth-sensitive). Think of it like a restaurant kitchen where the prep station and the pass are optimized for completely different workflows — most chips pick one. According to OpenAI, Jalapeño handles both by keeping model state closer to compute resources and integrating networking to cut chip-to-chip overhead.
“Serve more AI work per unit of power, while also returning responses more quickly.” — Richard Ho, OpenAI hardware head, via TechCrunch
Trust the Numbers – But Verify
Strong benchmark claims from a single source deserve scrutiny before reshaping any procurement strategy.
SemiAnalysis’ InferenceX benchmark is public, which matters. The results, however, are entirely OpenAI-reported — no independent third-party replication yet, as The Register and others have noted. Semiconductor analyst coverage echoes this caution: self-reported first-gen results are a credible opening argument, not a settled conclusion. It is also a meaningful caveat for infrastructure buyers making real purchasing decisions.
Until independent labs replicate these results, treat the numbers as a strong opening argument, not a verdict.
The stakes are concrete. OpenAI plans to deploy Jalapeño in its own infrastructure by end-2026, with second- and third-generation designs already underway. The chip went from initial design to tapeout in nine months — a timeline that signals how AI-assisted design is compressing hardware development cycles in ways that should unsettle traditional semiconductor timelines.
OpenAI is no longer purely a software company renting compute from Nvidia. Jalapeño is a strategic move toward owning the full stack — model, chip, and infrastructure. Whether these numbers survive independent scrutiny will determine if this is a legitimate threat to Nvidia’s dominance or an extraordinarily expensive press release.





























