Jalapeño Is OpenAI’s First Custom Chip: It Claims to Beat Nvidia With 1.9x More Efficiency

Broadcom-built inference chip claims up to 4x latency gains over Nvidia Blackwell, with full deployment targeted for end of 2026

Annemarije de Boer Avatar
Annemarije de Boer Avatar

By

Image: OpenAI

Key Takeaways

Key Takeaways

  • Jalapeño claims 1.5x–4.1x performance gains over Nvidia Blackwell in OpenAI’s own benchmarks.
  • Designed by OpenAI and Broadcom, Jalapeño optimizes both prefill and decode inference phases simultaneously.
  • OpenAI plans full Jalapeño infrastructure deployment by end-2026, with next-gen designs already underway.

Almost every ChatGPT response delivered to millions of users runs on Nvidia custom silicon — at serious power cost and serious expense to OpenAI. That bill has been accumulating for years. Now OpenAI has built its own chip: Jalapeño, developed with Broadcom, designed specifically for inference (serving model responses to users, not building the model). The early benchmark numbers are striking. The catch: OpenAI ran the tests themselves.

What the Numbers Actually Say

The reported gains are large enough that infrastructure buyers should pay close attention — with one important asterisk.

Across SemiAnalysis’ public InferenceX benchmark, testing three models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — OpenAI reported 1.5x to 1.9x more AI work per watt at peak throughput versus comparison systems, and 1.7x to 3.6x lower end-to-end latency. For interactive workloads, where a user is actively waiting for a response, performance reportedly jumped 2.1x to 4.1x higher. These are not rounding errors.

A few supporting details worth holding:

  • Chip rated at 700W; sustained measured power stayed at or below 550W during tested workloads
  • Models tested span 120B to 1 trillion parameters — a serious range
  • Broadcom CEO Hock Tan positioned it as comparable to Nvidia Blackwell and Google TPUs, according to Reuters
  • OpenAI used Codex and GPT-Astra to shorten design and verification cycles — AI helping build AI hardware

The architecture insight matters here. Inference runs in two phases: prefill (compute-heavy) and decode (memory-bandwidth-sensitive). Think of it like a restaurant kitchen where the prep station and the pass are optimized for completely different workflows — most chips pick one. According to OpenAI, Jalapeño handles both by keeping model state closer to compute resources and integrating networking to cut chip-to-chip overhead.

“Serve more AI work per unit of power, while also returning responses more quickly.” — Richard Ho, OpenAI hardware head, via TechCrunch

Trust the Numbers – But Verify

Strong benchmark claims from a single source deserve scrutiny before reshaping any procurement strategy.

SemiAnalysis’ InferenceX benchmark is public, which matters. The results, however, are entirely OpenAI-reported — no independent third-party replication yet, as The Register and others have noted. Semiconductor analyst coverage echoes this caution: self-reported first-gen results are a credible opening argument, not a settled conclusion. It is also a meaningful caveat for infrastructure buyers making real purchasing decisions.

Until independent labs replicate these results, treat the numbers as a strong opening argument, not a verdict.

The stakes are concrete. OpenAI plans to deploy Jalapeño in its own infrastructure by end-2026, with second- and third-generation designs already underway. The chip went from initial design to tapeout in nine months — a timeline that signals how AI-assisted design is compressing hardware development cycles in ways that should unsettle traditional semiconductor timelines.

OpenAI is no longer purely a software company renting compute from Nvidia. Jalapeño is a strategic move toward owning the full stack — model, chip, and infrastructure. Whether these numbers survive independent scrutiny will determine if this is a legitimate threat to Nvidia’s dominance or an extraordinarily expensive press release.

Share this

At Gadget Review, our guides, reviews, and news are driven by thorough human expertise and use our Trust Rating system and the True Score. AI assists in refining our editorial process, ensuring that every article is engaging, clear and succinct. See how we write our content here →