OpenAI Jalapeno Hot Chips Results: Up to 1.9x Efficiency
OpenAI shared first measured Jalapeno benchmarks at Hot Chips: up to 1.9x performance per watt and 3.6x lower latency vs. a Blackwell system.
OpenAI shared first measured Jalapeno benchmarks at Hot Chips: up to 1.9x performance per watt and 3.6x lower latency vs. a Blackwell system.
Introduction
On June 24, 2026, OpenAI announced Jalapeno, its first custom AI inference chip co-developed with Broadcom. Two months later, on August 25, 2026, the company disclosed the chip's first measured benchmark results in a talk by Richard Ho, OpenAI's head of hardware, at the Hot Chips conference. The talk attached measured numbers to design claims that had, until then, been stated only as targets.
Feature Overview
InferenceX Benchmark Methodology
OpenAI tested Jalapeno using InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request rather than isolated compute throughput. Three open models were run: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Results were normalized against each accelerator's published chip power rating, comparing Jalapeno to an Nvidia Blackwell system.
Aggregate Results Across Three Models
Across the three benchmarked models, OpenAI reported that Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput, and 1.7 to 3.6 times lower end-to-end latency than the comparison Blackwell system. For highly interactive workloads specifically, OpenAI reported 2.1 to 4.1 times higher performance. The company frames its evaluation standard as a "matched user experience" — work per unit of power while still meeting required latency — arguing this is a more useful yardstick than raw per-chip throughput.
Kimi K2.5 1T Results
On Kimi K2.5 1T, the largest public model in the test set at 1 trillion parameters, Jalapeno showed approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. OpenAI says these results place Jalapeno "on the Pareto frontier" across the tested operating range, meaning no comparison configuration beat it on both power efficiency and latency at the same time.
Power Characteristics
Jalapeno is rated at 700 watts, but OpenAI reports that measured sustained power on the tested workloads stayed at or below 550 watts — a gap between the chip's rated ceiling and its observed draw under the benchmarked conditions.
Architecture for Prefill and Decode
OpenAI describes Jalapeno's design as targeting both phases of inference: the compute-bound prefill phase and the memory-bandwidth-bound decode phase, calling it a "balanced and fungible accelerator." Model state, including the KV cache, "can be explicitly placed and kept local," and the network is described as "integral to the architecture" rather than a bolt-on interconnect, with a large domain intended to keep an entire workload inside one connected system. OpenAI also states its own models were used to help design and bring up the chip, and now assist with optimizing and programming it.
Usability Analysis
These are the first numbers to test OpenAI's original design claims for Jalapeno against actual measured performance rather than architectural projections. That distinction matters: the June announcement described a purpose-built inference ASIC and a fast development cycle, but offered no independent performance data. The Hot Chips figures, spanning three real open-weight models up to 1 trillion parameters, begin to fill that gap.
The choice of InferenceX as the benchmark is notable because it scores full request-serving behavior instead of isolated matrix-multiply throughput, which is closer to what an inference customer actually experiences. Still, the results come from OpenAI's own presentation and testing setup, using a comparison system's published power rating rather than its measured power draw — a normalization choice that favors interpretability but is not the same as an independently audited benchmark.
OpenAI calls Jalapeno "working first-party silicon with measured results" and frames it as "the beginning of a multigenerational platform," language that positions this Hot Chips talk as a milestone rather than a finished product story. Ho indicated deployment will begin at the end of 2026 in very small volumes, with more significant deployment in 2027.
Pros and Cons
Pros
- First measured, not just projected, performance data for Jalapeno across three real open-weight models
- Reports up to 1.9x performance per watt and up to 3.6x lower end-to-end latency versus the Blackwell comparison system
- Kimi K2.5 1T results demonstrate the architecture holds up on a trillion-parameter model
- Measured sustained power (at or below 550W) came in under the chip's 700W rating on tested workloads
Cons
- The comparison was against an Nvidia Blackwell system; Ho acknowledged competing hardware will likely advance significantly before Jalapeno's deployment
- Deployment begins at the end of 2026 in only "very small volumes," with more significant deployment not expected until 2027
- Power normalization used the comparison system's published power rating rather than its measured power draw
- Results are self-reported by OpenAI and have not yet been independently reproduced
Outlook
Ho's presentation converts Jalapeno from an announced roadmap item into a chip with disclosed performance figures, which is what any custom-silicon program needs before it can be evaluated on its own merits rather than its pitch. The next milestone to watch is the promised small-volume deployment at the end of 2026, followed by the more significant 2027 rollout Ho referenced — both of which will offer the first opportunity for real production data, as opposed to a controlled benchmark presentation. Given OpenAI's own acknowledgment that competing accelerators will keep advancing in the interim, Jalapeno's actual competitive position is likely to be reassessed once deployment volumes grow and third parties can compare it directly against whatever Nvidia and other custom-silicon programs ship by then.
Conclusion
The Hot Chips 2026 results give Jalapeno its first measured performance profile: up to 1.9x performance per watt and up to 3.6x lower latency than a comparison Nvidia Blackwell system across three open models, with a specific 1.5x-per-watt, 3.4x-latency figure on the trillion-parameter Kimi K2.5 model. The numbers are self-reported and the comparison methodology carries real caveats, but they represent a concrete step beyond the June announcement's design-only claims. This update is most relevant to infrastructure engineers and analysts already tracking OpenAI's custom-silicon program; the practical test of these figures will come with the small-volume deployment OpenAI has set for the end of 2026.
Editor's Verdict
OpenAI Jalapeno Hot Chips Results: Up to 1.9x Efficiency earns a solid recommendation within the GPT space.
The strongest case for paying attention: first measured, not just projected, performance data for Jalapeno across three real open-weight models. That alone raises the bar for what readers should expect in this space. Reinforcing that, reports up to 1.9x performance per watt and up to 3.6x lower end-to-end latency versus the Blackwell comparison system — practical value rather than just headline appeal. The broader signal worth registering is straightforward: Hot Chips 2026 marks the first time OpenAI has released measured performance data for Jalapeno, following the design-only June 2026 announcement. On the other side of the ledger, one constraint is real rather than a marketing footnote: the comparison was against an Nvidia Blackwell system; Ho acknowledged competing hardware will likely advance significantly before Jalapeno's deployment. It should factor into any serious decision. Layered on top of that, deployment begins at the end of 2026 in only very small volumes, with more significant deployment not expected until 2027 — which narrows the set of teams for whom this is an obvious yes.
For ChatGPT power users, OpenAI API customers, and enterprise teams already running on the OpenAI stack, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.
Pros
- First measured, not just projected, performance data for Jalapeno across three real open-weight models
- Reports up to 1.9x performance per watt and up to 3.6x lower end-to-end latency versus the Blackwell comparison system
- Kimi K2.5 1T results demonstrate the architecture holds up on a trillion-parameter model
- Measured sustained power (at or below 550W) came in under the chip's 700W rating on tested workloads
Cons
- The comparison was against an Nvidia Blackwell system; Ho acknowledged competing hardware will likely advance significantly before Jalapeno's deployment
- Deployment begins at the end of 2026 in only very small volumes, with more significant deployment not expected until 2027
- Power normalization used the comparison system's published power rating rather than its measured power draw
- Results are self-reported by OpenAI and have not yet been independently reproduced
References
Comments0
Key Features
1. First measured Hot Chips 2026 benchmark results for Jalapeno, following the design-only June 2026 announcement 2. 1.5 to 1.9x more AI work per watt and 1.7 to 3.6x lower end-to-end latency across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T versus a comparison Nvidia Blackwell system 3. On Kimi K2.5 1T specifically: approximately 1.5x peak performance per watt and 3.4x lower end-to-end latency 4. Measured sustained power at or below 550W against Jalapeno's 700W rating on the tested workloads 5. Architecture targets both prefill and decode phases, with KV cache locality and a network integral to the design 6. OpenAI's own models were used to help design, bring up, and now optimize and program the chip
Key Insights
- Hot Chips 2026 marks the first time OpenAI has released measured performance data for Jalapeno, following the design-only June 2026 announcement
- The reported 1.5-to-1.9x performance-per-watt gain and 1.7-to-3.6x latency reduction span three real open models, not a single test case
- On Kimi K2.5 1T, the largest model tested, Jalapeno still delivered roughly 1.5x performance per watt and 3.4x lower latency, suggesting the design holds up at trillion-parameter scale
- Measured sustained power stayed at or below 550W against a 700W rating, indicating the chip ran under its rated power ceiling during the benchmarked workloads
- The architecture explicitly targets both the compute-bound prefill phase and the memory-bandwidth-bound decode phase, aiming for a balanced, fungible accelerator rather than one tuned for a single inference stage
- The benchmark ran against an Nvidia Blackwell comparison system, and OpenAI's head of hardware acknowledged competing hardware will likely advance significantly before Jalapeno reaches deployment
- Deployment remains gated to very small volumes at the end of 2026, with more significant deployment not expected until 2027
- Power normalization was based on the comparison system's published chip rating rather than its measured power draw, a methodology choice that affects how directly the efficiency figures can be compared
Was this review helpful?
Share
Related AI Reviews
OpenAI Previews Private Safety Processing, No Data Kept
OpenAI previews Private Safety Processing, an automated abuse-monitoring system that extends Zero Data Retention to enterprise and API customers.
ChatGPT for Teens: Age Prediction Auto-Enrolls Under-18s
OpenAI launched ChatGPT for Teens on Aug 18, 2026, auto-enrolling predicted under-18 users with stricter safety limits and content restrictions.
OpenAI Disbands Preparedness Team Ahead of Expected IPO
OpenAI disbanded its Preparedness team for catastrophic AI risk, reassigning duties as it streamlines operations ahead of an expected IPO.
OpenAI Ultrafast Preview: GPT-5.6 Sol Runs Up to 14x Faster
OpenAI previews Ultrafast, a Cerebras-powered GPT-5.6 Sol tier running up to 14x faster, in limited preview to select API customers.
