news

OpenAI’s Jalapeño Beat GB300 on Its Test. NVIDIA’s $63 Bear Case Needs Three More Things.

OpenAI's Broadcom-built Jalapeño beat GB300 on throughput per watt. Our $63.15 NVIDIA bear case still needs capex digestion, share loss and a de-rating.

Jalapeño is one leg of the $63 NVIDIA case, not the whole case

OpenAI-reported benchmark results and R40 model scenarios

MetricReported resultScope
Throughput per watt1.5–1.9× GB200/300Inference only
End-to-end latency1.7–3.6× lowerOpenAI-reported
Initial deploymentLate 2026Small volume
NVIDIA base case$245.07R40 model
NVIDIA bear case$63.15All bear inputs

Jalapeño results are vendor-reported InferenceX comparisons against the best available GB200/GB300 systems, not production cost-per-token disclosures. Fair values are R40 model outputs; the bear case also assumes a capex digestion, networking share loss, lower margins and valuation compression.

OpenAI reported its first detailed performance results for Jalapeño on August 25. Across InferenceX workloads including GPT-OSS 120B, DeepSeek R1 and Kimi K2.5, the company says its custom chip produced 1.5–1.9× more throughput per watt at peak and 1.7–3.6× lower end-to-end latency than the best comparable NVIDIA GB200 and GB300 systems available for the test.

That is a material result. It is also a vendor-submitted benchmark of engineering silicon, not a production cost-per-token disclosure. It establishes that a purpose-built ASIC can beat Blackwell on the inference workloads it was built to run. It does not establish how many chips OpenAI can deploy, what each useful token costs after the rack and network are included, or how Jalapeño performs at training.

The distinction matters because our NVIDIA model already has a bear case whose note says “custom ASICs take the inference workload.” That case is worth $63.15 a share, 74.2% below the $245.07 base case. Jalapeño is evidence for the phrase in quotation marks. It is not evidence for every other assumption required to get from $245 to $63.

What Jalapeño is

Jalapeño is OpenAI’s first custom “Intelligence Processor”: an inference ASIC designed by OpenAI and industrialised with Broadcom and Celestica. It runs a trained model — prompt processing, token generation and the movement of the model’s KV cache — rather than training that model from scratch.

That narrower job is the point. A general-purpose GPU must support a large range of models, numerical formats, training operations and software. OpenAI knows the kernels, memory traffic and serving patterns behind ChatGPT, Codex and its API. It can trade generality for utilisation: keep weights, KV cache and intermediate results close to the compute; reduce movement across the package and rack; and balance prefill and decode for interactive and agentic workloads.

OpenAI’s June announcement disclosed the partners, the inference-only purpose, the nine-month design-to-tape-out cycle and a target for initial deployment by the end of 2026. It did not publish a full datasheet. The detailed configuration below is therefore a mixture of company disclosure, SemiAnalysis lab reporting and estimates from package and die photographs — not one set of final production specifications.

Aspect Reported Jalapeño configuration
Process TSMC N3P compute die; separate N3E I/O chiplet
Compute die Near-reticle, about 840 mm²; approximately 25.5 × 33 mm
Peak compute 13.4 PFLOPS MXFP4 on the B0 stepping
Compute design Weight-stationary systolic array with smaller matrix shapes and large local SRAM
Memory Six HBM4 stacks, 216 GB and 15.4 TB/s per package
Power 700 W rated TDP; no more than about 550 W sustained in reported tests
Fabric 32 × 800G SerDes on the I/O chiplet; 24 local and eight global links
Host and network PCIe Gen 5 host link; Broadcom Tomahawk at platform level

The architecture is homogeneous: the published result does not depend on fixing separate chips permanently to prefill and decode. It uses a weight-stationary, TPU-like data path but supports smaller matrix shapes so varied transformer GEMMs do not leave most of a large array idle. Large on-chip SRAM and 216 GB of HBM4 at 15.4 TB/s are there for the same reason — inference becomes a memory-movement problem before peak FLOPS run out.

Reports describe 128 accelerators per rack. Multiplying the package figures gives approximately 1.72 exaflops of MXFP4 peak, 27.65 TB of HBM4 and 1.97 PB/s of aggregate memory bandwidth per rack. Those three rack totals are our arithmetic, not separately disclosed system specifications. The fabric is designed to extend a connected domain toward roughly 2,048 accelerators so more model state stays local rather than crossing a conventional scale-out boundary.

The B0 stepping used for the reported 13.4-PFLOPS figure is expected to improve performance per watt by about 25% from the A0 silicon used in much of the early work. That is still a target until volume systems are running. Engineering samples exist now; small-volume deployment is planned for late 2026, with the meaningful ramp in 2027.

What the benchmark proves — and what it does not

InferenceX measures the trade-off every inference operator has to make. Large batches maximise total tokens and minimise unit cost, but make each user wait. Small batches improve interactivity and latency, but complete less aggregate work. A useful comparison therefore traces a curve across operating points rather than printing one peak number.

Jalapeño’s reported advantage appears across that curve: 1.5–1.9× throughput per watt at the peak and 1.7–3.6× lower end-to-end latency against the best GB200/GB300 comparisons available at the time. SemiAnalysis described its performance per watt as industry-leading across operating points, with some single-token-prediction results also ahead of published Vera Rubin figures.

Three caveats travel with that sentence.

First, the result is inference only. Training needs flexibility across changing model architectures, high-precision accumulation, large collective operations and a mature debugging and compiler environment. NVIDIA’s CUDA ecosystem is most valuable where the workload changes and developers need the same platform to do many jobs. Jalapeño attacks the more repeatable serving job where specialisation pays most.

Second, the comparison is early. OpenAI has engineering silicon; NVIDIA ships commercial systems at scale. The Jalapeño results also largely avoid multi-token prediction and fixed prefill-decode disaggregation, which makes the architecture impressive but does not make every software and system configuration directly interchangeable. NVIDIA’s own kernels, inference runtimes and Rubin systems will keep moving while Jalapeño ramps.

Third, performance per watt is not cost per token. Yield on a near-reticle die, HBM4 pricing, packaging, network content, utilisation, uptime and software engineering all sit between a lab watt and an income statement. OpenAI has not disclosed a production rack price, total cost of ownership, deployed chip count or token cost. Nobody outside the companies can convert 1.9× into a revenue loss for NVIDIA yet.

Jalapeño is one leg of the $63 case

Our NVIDIA base case is $245.07 a share. It already assumes competition: the price of an NVL72-class rack falls 1% every quarter, Data Center compute EBITDA margin glides from 76% to 66%, and the business eventually reaches a physical ceiling rather than compounding forever.

The published bear case is $63.15, but “custom ASICs take inference” is only one clause in it. The scenario simultaneously assumes:

Jalapeño directly supports the pressure on inference price, volume and margin. It says nothing by itself about hyperscaler capex shrinking. Its platform uses Broadcom Tomahawk networking, which supports the merchant-Ethernet clause, but one OpenAI design does not prove NVIDIA loses scale-out or NVLink economics across the market. And a benchmark cannot choose the multiple or discount rate investors apply five years from now.

The sensitivities show how much compression is hidden inside the label:

Change from the $245.07 base Fair value Effect
Discount rate 10% → 13% $217.40 −11.3%
Rack-price drift −1% → −2% a quarter $218.14 −11.0%
Rack-price drift −1% → −3% a quarter $195.43 −20.3%
Exit revenue multiple 8× → 4× $150.78 −38.5%
Full published bear case $63.15 −74.2%

The rack-price rows are sensitivities, not estimates of Jalapeño’s impact. They isolate the most direct economic channel — faster erosion in the price of general-purpose compute — while holding NVIDIA’s modeled rack volumes intact. Even tripling the base rate of quarterly price erosion to 3% takes fair value to $195.43, not $63.15. The full bear case only appears when slower demand, lower margins, networking losses and a harsher valuation all arrive together.

At NVIDIA’s $208.48 close on August 24, the bear case is 69.7% below the share price and the base case is 17.6% above it. The $181.92 gap between the two scenarios is not the value of Jalapeño. It is the value of nearly every important NVIDIA risk breaking in the same direction.

Why the risk is still material

Rejecting the shortcut does not make the competitive signal small.

Inference is where deployed AI meets usage. Every ChatGPT response, API completion and Codex step consumes it, and agentic products can multiply token demand through long-running, multi-turn work. The workload is large, repeatable and sensitive to electricity and latency — exactly the category in which a custom ASIC can remove general-purpose overhead.

OpenAI is also not an isolated experiment. Google has spent years moving workloads onto TPUs; Amazon sells Trainium; Meta and other hyperscalers are building custom accelerators. Broadcom and Marvell exist as public-market beneficiaries of that trend. Jalapeño matters because one of the largest prospective buyers of inference compute has now designed the substitute around its own production workload.

That can reduce NVIDIA’s share of inference even while NVIDIA’s inference revenue grows. If total token demand doubles and custom silicon takes a quarter of the incremental work, the merchant-GPU pool still expands. The bear case requires the denominator to disappoint as the share falls: demand digestion plus substitution, not substitution alone.

The same result is constructive for our Broadcom model. Its AI semiconductor vertical is worth $183.36 of a $282.09 base case, or 65%. Jalapeño is the kind of custom XPU plus networking programme that line is designed to capture. But Broadcom and OpenAI have disclosed neither programme revenue nor deployed volume, so the model does not change on either side today.

The model stays where it is

There is no model edit here. NVIDIA’s base already prices gradual competition through falling rack prices, a ten-point compute-margin glide and a finite volume ceiling. Moving to the published bear case would require evidence that Jalapeño is one example of a market-wide displacement, that total AI infrastructure demand is also weakening, and that NVIDIA is losing networking content as its systems mix changes.

The first InferenceX result is evidence about capability. The model needs evidence about deployment and economics: chips shipped, megawatts installed, cost per token in production, NVIDIA revenue displaced and the margin at which Broadcom supplies the replacement.

NVIDIA reports its July quarter on August 26. As our preview argues, the October revenue guide and Rubin pricing matter more to the near-term model than one engineering benchmark. A guide near $104 billion alongside stable margin would be direct evidence that absolute demand is still outrunning custom-silicon share loss. A guide miss with pricing pressure would begin to connect the benchmark to the bear case.

What to watch

  1. Deployed Jalapeño volume in 2027. Engineering samples and a late-2026 small-volume launch do not displace an installed base. A rack count, accelerator count or megawatt figure would.
  2. Production cost per million tokens. Performance per watt is one input. A disclosed all-in figure at matched latency would include yield, HBM4, networking, utilisation and software.
  3. The benchmark against Rubin at matched software. GB200 and GB300 are the commercial comparison today. Vera Rubin is the relevant generation when Jalapeño reaches volume.
  4. NVIDIA’s inference mix and rack pricing. The model assumes a $3.0 million NVL72-class rack and 1% quarterly price erosion. Faster erosion without a volume offset is the cleanest path from Jalapeño into fair value.
  5. Broadcom programme revenue. A named OpenAI contribution, deployment schedule or capacity commitment would show whether this is a successful product or only successful silicon.

A real threat is not the whole bear case

Jalapeño weakens the idea that every valuable AI token must run on NVIDIA. OpenAI has designed an inference chip around real transformer serving, brought engineering silicon up at target frequency and power, and reported a substantial advantage over Blackwell on the metrics inference buyers care about. That is not a demo to wave away.

But $63.15 is the price of a bundle: custom ASIC displacement, a capex digestion, merchant networking taking share, lower margins, slower terminal growth, a doubled valuation haircut and a higher discount rate. Jalapeño supplies evidence for the first item and a clue toward the third. It does not supply the bundle.


OpenAI’s June 24 announcement is the source for Jalapeño’s inference purpose, nine-month tape-out, engineering-sample status, OpenAI/Broadcom/Celestica roles and initial deployment target by year-end 2026. The 1.5–1.9× throughput-per-watt and 1.7–3.6× latency results are OpenAI-reported InferenceX results covered by SemiAnalysis, not R40 tests; the production cost and volume remain undisclosed. The node, I/O, compute, memory, power and fabric specifications are industry reporting based on lab access and package analysis; OpenAI has not published an exhaustive datasheet, and the ~840 mm² die size is a photograph-based estimate. Rack totals are R40 multiplication of the reported per-package figures. The $208.48 price is the August 24 close and will not be the number shown today. The $245.07 base, $63.15 bear, scenario assumptions, $183.36 Broadcom AI-semiconductor contribution and every sensitivity are R40 model outputs — our assumptions, not company guidance.

Related

Stocks in this article