AI
“OpenAI published measured performance results showing its custom AI inference chip Jalapeño outperforms Nvidia's GB300 on key inference metrics, delivering 1.5-1.9x more AI work per watt at peak throughput, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance for interactive workloads like AI agents, per the InferenceX…”
Plain restatementOpenAI has published first measured results for its Jalapeño inference chip, reporting three specific performance ratios versus Nvidia comparison systems, obtained using SemiAnalysis's public InferenceX benchmark.
Distortion codes this site does not recognise yet: benchmark_cherry_picking, scale_conflation, harness_mismatch, cost_compute_omission, misattribution. Not collectible until the field guide has an entry.
This one is mostly accurate. OpenAI really did publish these exact numbers for its Jalapeño inference chip on 25 August 2026, and the figures quoted in the post, 1.5 to 1.9 times more work per watt, 1.7 to 3.6 times lower latency, and 2.1 to 4.1 times better on highly interactive workloads, match OpenAI's own page word for word. InferenceX is a real public benchmark run by the research firm SemiAnalysis. The important missing context is that OpenAI ran the tests and supplied the numbers itself. SemiAnalysis watched some runs in OpenAI's lab and broadly backed the conclusion, but said it did not run the full benchmark suite, and it called the comparison against Nvidia's GB300 "somewhat incomplete and unfair" because the fair rival is Nvidia's newer Rubin platform, which was not tested. Two smaller points: the top of the work-per-watt range was measured against an older GB200, not a GB300, and the chip is still at engineering-sample stage with deployment only planned by the end of 2026. Also, the post's graphic credits the claim to Sam Altman, but it was published by OpenAI and presented by its hardware chief Richard Ho.
OpenAI published measured [drifted from the evidence:] performance results [drifted from the evidence:] showing its [drifted from the evidence:] custom AI inference chip Jalapeño [drifted from the evidence:] outperforms Nvidia's GB300 on key inference [drifted from the evidence:] metrics, delivering 1.5-1.9x more AI work per watt at peak throughput, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance [drifted from the evidence:] for interactive workloads like AI agents, per the InferenceX [drifted from the evidence:] SemiAnalysis benchmark
OpenAI [added by the neutral restatement:] has published [added by the neutral restatement:] first measured results [added by the neutral restatement:] for its Jalapeño inference [added by the neutral restatement:] chip, reporting three specific performance [added by the neutral restatement:] ratios versus Nvidia comparison systems, obtained using SemiAnalysis's public InferenceX benchmark.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- All three headline figures match OpenAI's published post verbatim: 1.5-1.9x work per watt at peak throughput, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x on highly interactive workloads.
- OpenAI did publish measured results from working silicon, presented at Hot Chips 2026 by hardware VP Richard Ho, and Jalapeño is a real chip co-developed with Broadcom, unveiled in June 2026.
- InferenceX is a real, public, open-source SemiAnalysis inference benchmark, not an invented artifact.
- GB300 is a genuine comparator in the published tables, and OpenAI's chip chief said GB300 was the leading option on the benchmark used.
- The three models named in the post's caption, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, are the models actually tested.
- The claimed deployment timing, beginning by end of 2026, matches OpenAI's stated plan.
- SemiAnalysis, the benchmark's operator, independently endorsed the broad conclusion, including that Jalapeño's figures exceed published Vera Rubin numbers on throughput per megawatt.
- Marketing as evidence: the phrase "per the InferenceX SemiAnalysis benchmark" reads to a general audience as third-party measurement. The benchmark is third-party; the runs and numbers are not. SemiAnalysis states plainly that OpenAI supplied all numbers, that it verified runs in person but did not execute the full suite, and that it had not seen AgentX results. The claim's own opening, "OpenAI published," partly offsets this, which is why this is a framing gap rather than a fabrication.
- Benchmark cherry picking: the comparator was selected favourably. GB300 is an HBM3E Blackwell Ultra part; Jalapeño uses HBM4. The benchmark operator itself called the Blackwell comparison incomplete and unfair and named Vera Rubin, which is shipping to customers now, as the like-for-like peer. Vera Rubin was not tested.
- Scale conflation: the claim attributes the full 1.5-1.9x range to GB300. OpenAI's own chart captions show the GPT-OSS-120B comparison, which anchors the top of that range, was run against a 1,200 W GB200, a prior-generation part, not a GB300.
- Harness mismatch: Jalapeño's single-token-prediction results were set against Nvidia STP configurations, while production Nvidia deployments commonly use multi-token prediction. SemiAnalysis's Rubin comparison likewise sets Jalapeño STP against Rubin MTP figures. Also, the tested scenario is nominal 8k/1k single-turn, not the long-context multi-turn AgentX scenario the operator considers most representative of agentic production load, which is notable given the claim invokes "AI agents."
- Cost compute omission: all per-watt ratios are normalised to rated package TDP rather than measured draw. OpenAI disclosed that Jalapeño's sustained power stayed at or below 550 W. The normalisation choice is disclosed by OpenAI but disappears entirely from the social-post version.
- Omitted qualifier: Jalapeño is at engineering-sample stage and is inference-only. It cannot train models, the workload where Nvidia is unchallenged, and it is not yet deployed.
- Misattribution (in the post's image, not the claim text): the graphic headline reads that "CEO Sam Altman claims it outperforms Nvidia's GB300 in key tests." The results were published by OpenAI and presented by hardware VP Richard Ho. Altman's contribution on X was a short remark that the chip is fast, not the GB300 comparison.
- No fully independent, end-to-end reproduction of the full InferenceX suite on Jalapeño exists as of 2026-08-28. SemiAnalysis's witnessing is meaningful corroboration but is explicitly not a full independent run.
- AgentX results, the operator's preferred agentic scenario, have not been published for Jalapeño. This bears directly on the claim's "interactive workloads like AI agents" framing.
- Per-model harness details beyond the chart captions, such as batch and concurrency sweeps behind the 2.1-4.1x interactive figure, were not retrievable in full.
- I read the OpenAI results page, the SemiAnalysis analysis and the InferenceX site through search excerpts rather than full page loads. The excerpts contain the decisive verbatim sentences, but I did not review the complete documents.
- Whether shipped Jalapeño hardware sustains these ratios at scale, against Vera Rubin, is untested and unknowable now.
The claim's numbers are reproduced accurately from OpenAI's own published post. OpenAI states that across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, "Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems," and that for highly interactive workloads it delivered 2.1 to 4.1 times higher performance. On Kimi K2.5, the largest public model tested, OpenAI reports roughly 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. The benchmark named is real. InferenceX, formerly InferenceMAX, is SemiAnalysis's open-source continuous inference benchmark research platform covering GB200 NVL72, GB300 NVL72, B200, MI355X and other accelerators. The decisive caveat comes from the benchmark operator itself. SemiAnalysis wrote that all numbers were provided to them by OpenAI, that they verified the InferenceX runs in person in the lab but did not run the full suite of InferenceX benchmarks and had not seen AgentX results, AgentX being their preferred suite for chip comparison because its long-context multi-turn characteristics reflect realistic production cache behaviour. SemiAnalysis further stated that the comparison to Blackwell is "somewhat incomplete and unfair" because Jalapeño competes against chips like Rubin that also use HBM4, that Vera Rubin systems are shipping to customers now while OpenAI has nothing beyond engineering samples, and that the models tested are not on the open frontier. Separately, however, SemiAnalysis concluded that even compared with Rubin, Jalapeño's single-token-prediction output throughput per megawatt surpasses the Vera Rubin multi-token-prediction figures Nvidia and CoreWeave published in July, and described the part as "beating every Nvidia, AMD, and Google chip we have been able to test". Comparator detail matters. OpenAI's own chart captions show two different Nvidia comparators: GPT-OSS-120B at nominal 8k/1k, STP, against a GB200 at 1,200 W package TDP, and DeepSeek R1 and Kimi K2.5 against a GB300 at 1,400 W package TDP, with Jalapeño at 700 W. Tom's Hardware notes that Jalapeño was not tested against Vera Rubin, does not train models, and that the major comparisons ran Jalapeño's single-token prediction against GB300 configurations doing the same even though Nvidia deployments commonly use multi-token prediction in production. Power normalisation used rated TDP, not measured draw. OpenAI revealed at Hot Chips that the processor is rated at 700 watts but measured sustained power stayed at or below 550 W on the workloads tested, and OpenAI normalised the benchmark results against each accelerator's published package TDP. Deployment status: OpenAI's first custom AI chip is expected to begin deployment in the company's computing infrastructure by the end of the year. Nvidia's response was dismissive rather than a technical rebuttal: Jensen Huang brushed off the chip, saying he remains confident in Nvidia's technology and ability to supply the AI industry even as customers develop their own processors.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/7a994b45d177/Z10hJ8V6-hEzZ3EqW_JoS7WUgK-
Ask this case
Answers come only from the case file above; nothing is added.
Did OpenAI actually publish these performance numbers for Jalapeño?
Yes. OpenAI published the exact figures on 25 August 2026: 1.5 to 1.9 times more AI work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance for interactive workloads. These match OpenAI's own post word for word.
Was this benchmark run independently by a third party?
Not fully. InferenceX is a real public benchmark from SemiAnalysis, but SemiAnalysis said OpenAI supplied all the numbers. SemiAnalysis verified some runs in person in OpenAI's lab but did not run the full benchmark suite and had not seen AgentX results, its preferred suite for chip comparisons.
Is comparing Jalapeño to Nvidia's GB300 a fair comparison?
SemiAnalysis itself called the GB300 comparison somewhat incomplete and unfair, saying the more appropriate rival is Nvidia's newer Rubin platform, which uses the same HBM4 memory as Jalapeño and is already shipping to customers. Rubin was not tested in this benchmark.
Is Jalapeño actually deployed and running AI workloads today?
No. Jalapeño is still at the engineering-sample stage. OpenAI's stated plan is to begin deployment in its computing infrastructure by the end of 2026, and the chip does not train models, only performs inference.
Were all the reported ratios measured against the same Nvidia chip?
No. OpenAI's own chart captions show the GPT-OSS-120B test, which sets the top of the 1.5 to 1.9x work-per-watt range, was run against an older GB200 chip, not the GB300 used in the other comparisons.