AI News

OpenAI Jalapeño vs Nvidia GB300: What the 104× Benchmark Claim Really Means

Evidence-first verdict: OpenAI’s Jalapeño results are genuinely impressive, but the viral interpretation is wrong. The chip was not shown to be 104× faster than Nvidia’s GB300. OpenAI reported 104.3× more mixed-token throughput per nominal package kilowatt at one matched low-latency operating point on DeepSeek R1. Across the more representative peak-throughput comparison, Jalapeño’s disclosed advantage was 1.5× to 1.9× in mixed throughput per kilowatt. That is still a consequential first result for custom silicon—but it is vendor-run, narrowly scoped and not independently reproduced.

OpenAI has published the first performance data for Jalapeño, its custom inference accelerator. The eye-catching result is a claimed 104.3× advantage over an Nvidia GB300 system. It is also the easiest result to misread.

The number does not compare maximum chip speed, total rack throughput or model quality. It compares how much mixed input-and-output token work each system delivered per kilowatt while both were constrained to Nvidia’s best disclosed per-user token cadence on one 8,000-input/1,000-output-token DeepSeek R1 test. Jalapeño sustained much more concurrent work at that particular latency target.

The stronger, less sensational conclusion is that Jalapeño occupied a better throughput-versus-latency curve in the three tests OpenAI disclosed. At peak throughput, it delivered 1.5× to 1.9× more mixed tokens per second per kilowatt; at the lowest-latency points, individual users received tokens 2.7× to 4.1× faster. Those results deserve attention. They do not yet establish production economics, broad workload leadership or superiority to Nvidia’s current Rubin generation.

Disclosure: Kingy.ai did not run Jalapeño or GB300 hardware. This audit uses OpenAI’s results and appendix, the public InferenceX methodology, SemiAnalysis’s account of what it observed, and Nvidia’s official platform specifications. Evidence labels follow Kingy.ai’s source and editorial standards. Research cutoff: August 25, 2026.

The result in 30 seconds

  • True: OpenAI reports a 104.3× matched-interactivity throughput-per-kilowatt result against GB300 on DeepSeek R1.
  • False: Jalapeño is not shown to be 104× faster than GB300 overall.
  • Most useful headline metric: At peak throughput, Jalapeño reports a 1.5× to 1.9× efficiency advantage across three selected workloads.
  • Most technically interesting result: Jalapeño combines high throughput with much lower time between output tokens, instead of winning only through large batches.
  • Evidence level: OpenAI supplied all numbers. SemiAnalysis says it witnessed selected InferenceX runs in OpenAI’s lab, but did not run the full suite and has not seen AgentX results.
  • Commercial reality: Jalapeño is an internal inference chip still undergoing production qualification. GB300 is a general-purpose training-and-inference platform available through Nvidia’s sales channel.

Complete Jalapeño benchmark table

OpenAI disclosed four comparisons for each of three models. The GPT-OSS test used GB200; the DeepSeek and Kimi tests used GB300. All were InferenceX, nominal 8k/1k, single-token prediction (STP) runs. OpenAI normalized throughput with package power ratings of 700 W for Jalapeño, 1,200 W for GB200 and 1,400 W for GB300.

All benchmark values disclosed by OpenAI on August 25, 2026
Workload Comparison Metric Jalapeño Nvidia Reported advantage
GPT-OSS 120B GB200 Peak mixed TPS/kW 85,448 44,960 ≈1.9× higher
GPT-OSS 120B GB200 End-to-end latency 1.03 s 1.80 s ≈1.7× lower
GPT-OSS 120B GB200 Minimum TBT 0.69 ms
(1,459 tok/s/user)
1.87 ms
(535 tok/s/user)
≈2.7× lower TBT
GPT-OSS 120B GB200 Mixed TPS/kW at Nvidia’s minimum TBT 22,935 427 ≈53.7× higher
at 535.28 tok/s/user
DeepSeek R1 670B, MXFP4 GB300 Peak mixed TPS/kW 19,641 11,781 ≈1.7× higher
DeepSeek R1 670B, MXFP4 GB300 End-to-end latency 1.65 s 5.99 s ≈3.6× lower
DeepSeek R1 670B, MXFP4 GB300 Minimum TBT 1.43 ms
(700 tok/s/user)
5.90 ms
(169 tok/s/user)
≈4.1× lower TBT
DeepSeek R1 670B, MXFP4 GB300 Mixed TPS/kW at Nvidia’s minimum TBT 12,258 118 ≈104.3× higher
at 169.41 tok/s/user
Kimi K2.5 1T, MXFP4 GB300 Peak mixed TPS/kW 18,195 11,862 ≈1.5× higher
Kimi K2.5 1T, MXFP4 GB300 End-to-end latency 1.56 s 5.31 s ≈3.4× lower
Kimi K2.5 1T, MXFP4 GB300 Minimum TBT 1.44 ms
(694 tok/s/user)
5.48 ms
(182 tok/s/user)
≈3.8× lower TBT
Kimi K2.5 1T, MXFP4 GB300 Mixed TPS/kW at Nvidia’s minimum TBT 6,744 120 ≈56.1× higher
at 182.46 tok/s/user

Source: OpenAI, “Jalapeño’s first results show industry-leading speed and efficiency in AI inference”. Ratios are OpenAI’s reported figures. Some will differ slightly if calculated from the displayed values because those values are rounded. For example, 12,258 ÷ 118 is about 103.9, while OpenAI reports 104.3× from the underlying data.

What the 104.3× claim actually measures

Imagine two restaurants promising that each diner receives one course every six minutes. One kitchen can serve only a nearly empty dining room at that pace; the other can maintain the same pace while serving many tables. The second kitchen is not making each diner’s meal 100 times faster. It is sustaining far more total work without making any diner wait longer between courses.

That is the shape of OpenAI’s 104.3× comparison.

On DeepSeek R1, GB300’s lowest reported time between tokens was 5.90 milliseconds, equivalent to roughly 169 output tokens per second for one user. At that same 169.41-token-per-second target, OpenAI reports:

  • Jalapeño: 12,258 mixed TPS/kW
  • GB300: 118 mixed TPS/kW
  • OpenAI’s ratio from unrounded data: 104.3×

The ratio becomes enormous because it uses GB300’s minimum-latency edge of the operating curve, where concurrency and total throughput are low. Jalapeño could apparently accept much more aggregate load while preserving that token cadence. This is a meaningful serving result for interactive systems. It is not a maximum-throughput-versus-maximum-throughput comparison.

At each platform’s peak-throughput point on the same DeepSeek test, the difference was 19,641 versus 11,781 mixed TPS/kW: about 1.7×, not 104×.

The methodology, translated into plain English

Mixed TPS/kW is an efficiency metric, not “words per second”

TPS means tokens per second. InferenceX’s public benchmark client defines total throughput from processed input tokens plus generated output tokens. OpenAI’s mixed TPS therefore should not be read as output-token speed alone. The public post does not provide separate prompt-token and generated-token totals for each point, and an 8k/1k workload contains roughly eight times as many nominal input tokens as output tokens.

Dividing by kilowatts rewards systems that complete more token processing for a given accelerator power rating. That is useful for constrained data centers. It is not the same as wall-socket efficiency or total cost of ownership.

TBT measures how quickly the response keeps arriving

Time between tokens (TBT) is the delay from one generated token to the next. Lower is better. A 5.90 ms TBT is about 169 tokens per second per user; 1.43 ms is about 699, which explains OpenAI’s rounded 700-token figure.

TBT is what makes a streamed response feel fast after generation begins. It does not tell you how long the user waited for the first token.

TTFT is missing from OpenAI’s summary table

Time to first token (TTFT) measures the wait before the response starts. InferenceX can record it, but OpenAI’s appendix does not publish TTFT for these comparisons. A system could have excellent token cadence after starting yet still make the user wait on prompt processing. The disclosed end-to-end numbers partly constrain that risk, but they do not replace the missing TTFT breakdown.

End-to-end latency is the whole request

This measures the total time to process the prompt and generate the requested output. OpenAI reports 1.7× to 3.6× lower end-to-end latency across the three workloads. The page does not expose enough request-level detail to reconstruct those values or determine whether each latency cell comes from the same configuration as its peak-throughput cell.

“Nominal 8k/1k” is a synthetic request shape

The benchmark sends roughly 8,000 input tokens and asks for roughly 1,000 output tokens. That is a legitimate fixed-sequence stress test, but it is only one traffic pattern. It does not reproduce long, multi-turn agent sessions, shared prefixes, pauses, subagents or repeated KV-cache reuse.

STP means one predicted token per forward pass

The public InferenceX configuration guide defines single-token prediction (STP) as ordinary autoregressive decoding with one token per forward pass. Multi-token prediction (MTP) or speculative decoding can propose several tokens in a pass. OpenAI’s table uses STP on both sides, which avoids attributing a speculative-decoding advantage to Jalapeño.

A Pareto frontier is a curve, not one magic setting

An inference server trades concurrency against latency. Large batches can raise total throughput while slowing each user’s stream. A point is on the Pareto frontier when no tested alternative is both faster for the user and more productive per kilowatt.

OpenAI says Jalapeño’s curve sat above and to the better-latency side of the Nvidia comparison curves. The peak, minimum-TBT and matched-TBT rows therefore describe different operating points. They should not be combined as if one configuration delivered all maxima simultaneously.

Jalapeño vs GB300: disclosed specifications and status

This is not a conventional product-card comparison. OpenAI has disclosed few official Jalapeño specifications, while GB300 is sold as a documented rack-scale platform. “Not disclosed” is more honest than filling the gaps with estimates.

Publicly documented comparison as of August 25, 2026
Field OpenAI Jalapeño Nvidia GB300 NVL72 / Blackwell Ultra
Primary role Custom inference accelerator for OpenAI Training and inference platform
Benchmark package power rating 700 W; OpenAI says sustained measured chip power stayed at or below 550 W in tested workloads 1,400 W per accelerator in OpenAI’s normalization
Rack configuration Not officially quantified in the benchmark post 72 Blackwell Ultra GPUs and 36 Grace CPUs
GPU memory Not officially disclosed in the benchmark post 20 TB total GPU memory; up to 576 TB/s aggregate bandwidth
Scale-up interconnect Integrated network and a large connected domain described; bandwidth and topology not disclosed 130 TB/s NVLink bandwidth
Low-precision compute DeepSeek R1 and Kimi K2.5 results labelled MXFP4; peak hardware FLOPS not disclosed by OpenAI 1,080 dense FP4 PFLOPS across NVL72; 1,440 PFLOPS with sparsity
Software maturity Production qualification and software maturation still underway Commercial platform with Nvidia software and ecosystem support
Availability Planned internal deployment by the end of 2026; not offered for customer purchase Nvidia directs customers to sales

Sources: OpenAI’s Jalapeño results and Nvidia’s official GB300 NVL72 specifications. The published benchmark does not fully document the serving topology behind each operating point, so the table does not infer a one-chip system.

What each workload tells us

GPT-OSS 120B: the largest peak-efficiency lead

Jalapeño’s best disclosed peak ratio appears on GPT-OSS 120B: 85,448 versus 44,960 mixed TPS/kW, or about 1.9×. It also reached a 0.69 ms minimum TBT, compared with 1.87 ms for GB200.

This is the cleanest evidence that Jalapeño can do more than win a low-concurrency latency contest. The caveat is that the comparison is to GB200, not GB300, so it should not be used as a direct GB300 claim.

DeepSeek R1 670B: the source of 104.3×

DeepSeek produced Jalapeño’s most dramatic matched-interactivity ratio and its largest minimum-TBT lead. The combination—about 1.7× higher peak mixed TPS/kW and 4.1× lower minimum TBT—suggests a serving architecture that retains useful utilization at low latency.

But the 104.3× result sits at a selected edge of the curve. It is better understood as evidence of latency headroom than a universal speed multiplier.

Kimi K2.5 1T: the largest disclosed model

On the one-trillion-parameter Kimi K2.5 workload, Jalapeño reports 1.5× higher peak mixed TPS/kW, 3.4× lower end-to-end latency and 3.8× lower minimum TBT than GB300. This is arguably more important than the headline ratio because the advantage persisted on the largest disclosed model.

The comparison still uses a fixed 8k/1k single-turn shape. It does not show how the systems behave when a real agent repeatedly extends a long context or fans out into parallel tasks.

Does performance per kilowatt mean lower AI costs?

Potentially—but the benchmark does not prove it.

Using published package ratings makes the comparison consistent: 700 W, 1,200 W and 1,400 W. OpenAI also says Jalapeño drew no more than 550 W of sustained chip power on the tested workloads, but it normalized Jalapeño at the full 700 W rating. That is conservative for Jalapeño inside this particular calculation.

The missing denominator is the rest of the system. CPUs, memory outside the accelerator package, network switches, power conversion, cooling and idle reserve all consume energy. OpenAI does not disclose matched measured wall power for the Nvidia systems or an all-in rack figure for Jalapeño. Package-TDP-normalized TPS/kW therefore cannot establish data-center performance per watt.

It also cannot establish cost per successful AI job. Acquisition cost, yield, software engineering, utilization, failure rates, support and model quality all affect the economic result. No public Jalapeño price exists because the chip is for OpenAI’s own infrastructure.

A useful arithmetic check: peak throughput per rated package

Multiplying OpenAI’s peak TPS/kW values by the stated package ratings reverses the normalization. This is a derived check—not a separately reported benchmark—and should be read as average normalized throughput per accelerator package, not proof that one package ran the complete model alone.

Workload Jalapeño derived mixed TPS/package Nvidia derived mixed TPS/package Relationship
GPT-OSS 120B 59,814 53,952 (GB200) Jalapeño ≈1.11×
DeepSeek R1 670B 13,749 16,493 (GB300) Jalapeño ≈0.83×
Kimi K2.5 1T 12,737 16,607 (GB300) Jalapeño ≈0.77×

That does not undermine the reported efficiency win; it explains it. On DeepSeek and Kimi, GB300’s derived peak mixed throughput per rated package is higher, while Jalapeño completes more work per rated watt because its denominator is half as large. Buyers should not translate the efficiency ratio into a raw per-package speed claim.

What was validated about model quality?

Faster wrong answers would be a useless victory, so quality parity matters. SemiAnalysis says it confirmed that Jalapeño’s GSM8K evaluations were on par with Nvidia for the tested models. That is a useful sanity check, not a complete quality audit.

InferenceX’s public evaluation documentation describes GSM8K exact-match testing through the LM Evaluation Harness. OpenAI’s article does not publish the exact scores, sample outputs, tolerances or run links. GSM8K also measures a narrow grade-school-math capability; it cannot establish parity across coding, reasoning, long-context retrieval, safety or output stability.

Verdict on evals: no disclosed sign that Jalapeño bought speed by visibly breaking GSM8K, but not enough public evidence to claim broad numerical or behavioral equivalence.

How strong is the evidence?

Evidence layer What is available What it supports What it does not support
OpenAI primary source Methods summary, curves and 12 result cells The vendor’s exact claims and test framing Independent reproduction
Third-party observation SemiAnalysis says it witnessed selected lab runs Greater confidence that working A0 silicon produced real runs A third party controlling the full experiment
Open benchmark framework InferenceX code, definitions and reproducibility process Auditable methodology in principle A Jalapeño run package linked from OpenAI’s post
Quality check SemiAnalysis reports GSM8K parity A narrow correctness sanity check Broad model-quality parity
Production evidence Deployment plan and qualification work OpenAI intends to operate the chip internally Fleet reliability, cost, yield or sustained production performance

The evidence is stronger than a simulation or a peak-FLOPS slide: there is working A0 engineering silicon, measured model-serving data and a named public framework. It is weaker than a reproducible neutral benchmark with raw artifacts and matched system telemetry.

For context, a documented 16-GPU GB300 serving deployment illustrates how much topology and software detail can matter to real inference results. The OpenAI summary does not expose that level of deployment detail for each Jalapeño point.

The limitations OpenAI’s chart cannot answer

  1. The vendor supplied the numbers. SemiAnalysis observed selected runs, but says it did not execute the full InferenceX suite itself.
  2. There is no independent reproduction. OpenAI’s post does not link a Jalapeño recipe, public workflow run, raw request data or telemetry bundle.
  3. Only one request shape is shown. All disclosed comparisons are nominal 8k/1k fixed-sequence tests.
  4. No AgentX result is public. InferenceX’s AgentX scenario exercises long, multi-turn sessions, shared prefixes, pauses, subagents and KV-cache reuse—the features most relevant to the agentic workloads OpenAI emphasizes.
  5. The systems are normalized by package ratings, not measured facility power. This cannot establish rack-level energy efficiency.
  6. Important configuration details are absent. The summary does not list exact serving-software versions, complete parallelism and topology, concurrency for every selected point, or confidence intervals.
  7. The cells represent multiple operating points. Peak throughput, minimum TBT and matched-TBT throughput are not simultaneous properties of one setting.
  8. Quality evidence is narrow. “GSM8K parity” is reported without public scores and is not a broad eval suite.
  9. Production economics are unknown. There is no public price, yield, uptime, fleet utilization, maintenance or cost-per-token result.
  10. Availability is asymmetric. Jalapeño is engineering silicon moving toward an internal deployment; GB300 is a commercial platform with a mature customer ecosystem.

Why Nvidia Rubin matters more than the headline suggests

GB300 is the comparison OpenAI published, so it is the comparison this article audits. It is not the last word on Nvidia.

SemiAnalysis argues that Rubin is the more contemporary peer because both Rubin and Jalapeño belong to the HBM4 era. Nvidia says Vera Rubin NVL72 is ramping into full production and lists 72 Rubin GPUs, 20.7 TB of HBM4, 1,580 TB/s of aggregate memory bandwidth and 260 TB/s of NVLink bandwidth. Nvidia marks the figures as preliminary and subject to change.

Rubin also supports training, while Jalapeño is designed for inference. A custom inference ASIC can omit capabilities that a merchant platform needs to serve many customers and workloads. That specialization may be exactly why Jalapeño is efficient for OpenAI, but it complicates any claim about the “better chip.”

There is no public, same-lab, matched-software Jalapeño-versus-Rubin InferenceX result. Until one exists, claims that Jalapeño beats Nvidia’s current generation are extrapolation. Nvidia is also broadening its low-latency stack, including its production Groq inference push, so the competitive comparison will not remain a two-chip contest.

What this means for OpenAI, Nvidia and users

For OpenAI, Jalapeño looks like a credible path to reduce dependence on merchant-GPU economics for high-volume inference. Owning the model, compiler, serving software, network and silicon can produce advantages that a general platform cannot optimize for one customer. The bigger strategic result is not 104×; it is that first-generation silicon appears competitive enough to justify a multigenerational roadmap.

For Nvidia, this is pressure at the hyperscaler margin, not evidence of collapse. Jalapeño is not sold to enterprises, does not replace training accelerators and is not yet operating as a proven production fleet. OpenAI explicitly says it will continue widely deploying Nvidia and other partner accelerators for training and inference.

For users, there is no demonstrated API price cut, availability change or product guarantee. Faster low-latency silicon could eventually make agents feel more responsive and increase capacity. Whether savings reach customers depends on production scale, demand and pricing decisions—not a benchmark chart. Readers looking for the broader chip program can use our full OpenAI Jalapeño guide; this article is intentionally limited to the benchmark evidence.

What would change this verdict?

The verdict would become substantially stronger with:

  • public Jalapeño InferenceX run IDs, recipes, logs, raw artifacts and power telemetry;
  • a neutral party controlling repeated runs on both platforms;
  • AgentX and additional fixed-sequence shapes such as 1k/8k and long-context tests;
  • exact TTFT, error-rate and percentile distributions rather than selected summary points;
  • matched measured wall power and complete serving topology;
  • published quality scores across math, coding, reasoning and long context;
  • production data covering uptime, yield, fleet utilization and cost per successful job;
  • a same-generation Jalapeño-versus-Rubin comparison.

Final verdict

Jalapeño has earned “promising breakthrough,” not “104× Nvidia killer.”

The 104.3× claim is a legitimate vendor-reported ratio for one carefully defined operating point: mixed-token throughput per nominal package kilowatt at GB300’s minimum time between tokens on DeepSeek R1. It demonstrates that Jalapeño can apparently serve far more concurrent work while preserving a highly interactive token cadence.

The result to remember is the less viral one. Across three selected workloads, OpenAI reports 1.5× to 1.9× higher peak mixed throughput per kilowatt, 1.7× to 3.6× lower end-to-end latency and 2.7× to 4.1× lower minimum TBT. That combination is technically significant for a first custom inference chip.

The remaining gap is evidence, not arithmetic. These are OpenAI-run engineering-sample results on a single synthetic request shape, without public AgentX data, neutral reproduction, full power accounting or production economics. Jalapeño may become a major inference platform inside OpenAI. This benchmark is a strong opening result, not the final ranking.

Frequently asked questions

Is OpenAI Jalapeño really 104× faster than Nvidia GB300?

No. OpenAI reports 104.3× more mixed-token throughput per nominal package kilowatt at one matched low-latency point on DeepSeek R1. At each system’s peak-throughput point, the reported Jalapeño advantage on that workload is about 1.7×.

What does “104.3× at previous TBT” mean?

OpenAI fixed both systems at GB300’s best reported token cadence—169.41 output tokens per second per user—then compared aggregate mixed-token throughput per kilowatt. Jalapeño reportedly sustained 12,258 mixed TPS/kW versus 118 for GB300.

Why do 12,258 and 118 not equal exactly 104.3?

The displayed values are rounded. They divide to about 103.9. OpenAI reports 104.3× from its underlying, higher-precision measurements.

Did Jalapeño beat GB300 at peak throughput?

On the two disclosed GB300 workloads, yes on normalized mixed throughput per kilowatt: about 1.7× for DeepSeek R1 and 1.5× for Kimi K2.5. The benchmark does not show that Jalapeño had higher absolute rack throughput or lower total facility power.

What is the difference between TBT and TTFT?

TBT is the delay between generated tokens after a response starts. TTFT is the initial wait before the first token appears. OpenAI publishes minimum TBT and end-to-end latency here, but not a TTFT table.

Were the benchmarks independently verified?

Not fully. SemiAnalysis says it witnessed selected InferenceX runs in OpenAI’s lab, but all numbers came from OpenAI, it did not run the complete suite and it has not seen AgentX results.

Did Jalapeño preserve model accuracy?

SemiAnalysis reports GSM8K results on par with Nvidia, but neither its public summary nor OpenAI’s article provides exact scores. That is a narrow correctness check, not proof of broad quality parity.

Can companies buy Jalapeño?

No public product or price has been announced. OpenAI plans to begin deploying Jalapeño in its own compute infrastructure by the end of 2026.

Does Jalapeño replace Nvidia GPUs at OpenAI?

No. OpenAI says it will continue widely deploying Nvidia and other partner accelerators for both training and inference. Jalapeño is an additional internal inference platform.

Is GB300 the fairest Nvidia comparison?

It is the latest Nvidia system in OpenAI’s published table, but Rubin is the more current HBM4-generation peer. No controlled Jalapeño-versus-Rubin InferenceX comparison is public.

Primary sources and methodology