Case TS-DDEDFA1E5 Sept 2026benchmarkCompound claim

AI

“OpenAI launched GPT-6 Astra, a new AI model achieving 98.6% on ARC-AGI-3, 100% on ExploitBench, 97.6% on FrontierMath Tier 4 (v2), and 74.1% on DeepSWE v1.1, outperforming GPT-5.6 Sol and Claude Fable 5.1”

Plain restatementOpenAI released a model named GPT-6 Astra. On four named evaluations it recorded the scores 98.6%, 100%, 97.6% and 74.1% respectively, and these results place it ahead of GPT-5.6 Sol and Claude Fable 5.1.

Source exists but framing is misleadingConfidence High
What this verdict means →

Distortion codes this site does not recognise yet: harness_mismatch, cost_compute_omission, eval_contamination, benchmark_cherry_picking, capability_extrapolation. Not collectible until the field guide has an entry.

OpenAI really did launch GPT-6 Astra on September 3, 2026, and all four numbers in this post are real rather than invented. The problem is what was left out. The 98.6% on ARC-AGI-3 came from a special OpenAI-supplied test setup; the benchmark's own operator, ARC Prize, ran the identical model through its neutral setup and got 62.7%, and reporting indicates the setup behind the high score is not part of the product people can buy. The 100% on ExploitBench came from a test OpenAI itself flagged as possibly contaminated, and on the clean version it built to correct for that, the model scored around 39%. The claim that Astra outperforms Claude Fable 5.1 is contradicted by OpenAI's own comparison table, where Fable 5.1 wins on Humanity's Last Exam, and by independent benchmarking firm Artificial Analysis, which rated Astra level with its own predecessor and behind Fable 5.1 while costing 2.5 times more. On the coding benchmark cited, Astra's 74.1% is statistically tied with two rival models and behind Meta's Muse Spark 1.3 at 75.4%. Astra does appear genuinely strong at computer use, cybersecurity and hard mathematics, but the sweeping framing goes well beyond what the evidence supports.

The drift / as claimed vs as evidenced

OpenAI [drifted from the evidence:] launched GPT-6 Astra, a [drifted from the evidence:] new AI model [drifted from the evidence:] achieving 98.6% on [drifted from the evidence:] ARC-AGI-3, 100% [drifted from the evidence:] on ExploitBench, 97.6% [drifted from the evidence:] on FrontierMath Tier 4 (v2), and 74.1% [drifted from the evidence:] on DeepSWE v1.1, outperforming GPT-5.6 Sol and Claude Fable 5.1


OpenAI [added by the neutral restatement:] released a model [added by the neutral restatement:] named GPT-6 Astra. On [added by the neutral restatement:] four named evaluations it recorded the scores 98.6%, 100%, 97.6% and 74.1% [added by the neutral restatement:] respectively, and these results place it ahead of GPT-5.6 Sol and Claude Fable 5.1.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
⌿ Omitted qualifier
A load-bearing condition from the source quietly disappears from the claim.
harness_mismatch
cost_compute_omission
eval_contamination
benchmark_cherry_picking
▲ Exaggeration
A real finding gets inflated: stronger, bigger, faster, or more certain than the evidence supports.
$ Marketing as evidence
Promotional material dressed up as independent proof.
capability_extrapolation
Secondary sourcenamed-outlet trade journalism
The New Stack, "OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3"
Secondary sourcenamed-outlet and vendor-adjacent analysis
VentureBeat, Fortune, DataCamp, Vellum, officechai benchmark write-ups
Secondary sourcecommentary, superseded by operator verification (see below)
CryptoBriefing, "Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don't hold up to scrutiny"
Primary sourcebenchmark operator of record, independent evaluator
ARC Prize, "GPT-6 Astra on ARC-AGI-3" blog and verified results page
Primary sourcevendor channel of record (duality rule applies)
OpenAI, "GPT-6 Astra: A new generation of intelligence" launch page including evaluation footnotes
Primary sourcevendor safety documentation
OpenAI, "Path to Astra: critical capabilities and frontier safeguards"
Primary sourcevendor system card
OpenAI, GPT-6 Astra System Card, Deployment Safety Hub
Primary sourcebenchmark operator
Epoch AI, FrontierMath Tier 4 (v2) benchmark page with conflict-of-interest statement
Primary sourceindependent benchmark operator
Datacurve DeepSWE leaderboard and v1.1 methodology page
Primary sourceindependent evaluator with published methodology
Artificial Analysis, "Benchmarking GPT-6 Astra," Intelligence Index v4.1.1
Primary sourcevendor documentation
OpenAI API docs, GPT-6 Astra model page
● Primary source found
What is true
  • OpenAI did launch GPT-6 Astra on September 3, 2026. This is confirmed on OpenAI's launch page, system card, API documentation and partner channels.
  • The staged rollout described in the caption matches OpenAI's own wording: limited organizations first, expanding over coming days to Plus, Pro, Business and Enterprise users, plus the API and AWS.
  • 98.6% on ARC-AGI-3 is a real, operator-verified result under a specified harness. It is not invented.
  • 100% on ExploitBench is stated by OpenAI in its own preparedness document.
  • 97.6% on FrontierMath Tier 4 v2 and 74.1% on DeepSWE v1.1 both appear in OpenAI's published table, and the DeepSWE figure is independently corroborated by Datacurve within rounding.
  • Astra does clearly lead on several axes: computer use, cybersecurity capability, hard mathematics, token efficiency, and long-horizon agentic work.
  • Greg Brockman's "Welcome to the AGI era" remark is confirmed by multiple outlets that attended the pre-launch briefing.
What is misleading
  • Omitted qualifier: The claim presents 98.6% on ARC-AGI-3 as a property of the model. The operator's own verified results show the same model scoring 62.7% on the same benchmark through ARC Prize's provider-neutral harness. The number is a property of the model plus a specific stateful scaffold, and reporting says that scaffold is not what customers buy. A reader takes away "the model solves ARC-AGI-3," which is not what was measured.
  • Cost compute omission: The headline runs cost roughly $19,000 to $26,000 each and used maximum or high reasoning effort. OpenAI states its tables report the maximum score at any effort. None of this survives into the claim, which implies a routine capability.
  • Eval contamination: The 100% on ExploitBench is presented as a clean sweep. OpenAI itself flagged contamination concerns on that benchmark and built a controlled replacement, on which Astra scores roughly 39%. OpenAI also notes 100% may not be achievable in principle under the evaluation's constraints. The claim propagates the compromised number and drops the corrective one, which is the single largest gap between claim and evidence here.
  • Benchmark cherry picking: On DeepSWE the claim uses Fable 5.1's 67.4% as the reference point. Opus 5 (73.7%), Gemini 3.8 Flash (73.8%) and Muse Spark 1.3 (75.4%) are all at or above Astra, and Muse Spark led the public leaderboard on the day the post ran. The lead is manufactured by comparator selection.
  • Margin presented beyond noise: A 0.4 point gap on a 113-task benchmark with a published ±3 band is not a ranking. It is a tie.
  • Exaggeration by overgeneralization: "Outperforming GPT-5.6 Sol and Claude Fable 5.1" is stated without scope. OpenAI's own table shows Fable 5.1 ahead on Humanity's Last Exam with tools, 65.0% to 57.2%, and Fable 5 and Opus 5 ahead on FrontierCode splits. Independently, Artificial Analysis places Astra level with Sol and behind Fable 5.1 overall, at 2.5 times the price. The sweeping form of the comparison is contradicted by the very table it was drawn from.
  • Marketing as evidence: Three of the four numbers are vendor-run and were reproduced as neutral fact. For quality and superiority claims, the vendor channel establishes what OpenAI reports, not what is independently true. The FrontierMath figure additionally rests on a benchmark whose operator publicly discloses OpenAI funding and exclusive access to part of the problem set.
  • Capability extrapolation: The caption's "Welcome to the AGI era" and "new highs for the industry across... coding" convert a mixed benchmark table into a general intelligence verdict. The coding claim in particular is contradicted by the DeepSWE and FrontierCode evidence.
What is uncertain
  • The original ExploitBench's total item count and composition are not established in what I retrieved. The 20-vulnerability figure belongs to OpenAI's contamination-controlled internal port, not necessarily the benchmark that produced the 100%.
  • No independent re-run of Astra on FrontierMath Tier 4 v2 was found. The 97.6% rests entirely on OpenAI's report, on a benchmark OpenAI funded.
  • The exact discrepancy between the post's 98.6% and OpenAI's headline 99.9% resolves to different reasoning-effort settings under the same harness, but I did not retrieve a statement from OpenAI explaining why the lower figure circulated first.
  • Whether ARC Prize's harness-labelling change alters how Astra's result is displayed going forward is announced but not yet observable.
  • Artificial Analysis index values are reported slightly differently across write-ups (61 versus 61.2, 65.7 versus 66). The ordering is consistent; the decimals are not load-bearing here.
Evidence summary

The release is real and fully documented on official channels. OpenAI published a launch page, a system card, a preparedness write-up and an API model page for GPT-6 Astra on September 3, 2026, and Microsoft published a Foundry availability post. This part of the claim needed no adjudication. All four numbers are real and traceable. None is fabricated. What the sources show is that each one carries a condition that the post removed. ARC-AGI-3. ARC Prize, the benchmark's operator, ran the model itself and published verified results. With ARC Prize's own provider-neutral Standard harness, Astra scored 62.7% at max reasoning, at a run cost near $26,000. With OpenAI's Provider Adapter harness, which ARC Prize describes as preserving opaque reasoning state between requests and using compaction, the model reached 99.9% at high reasoning. The 98.6% figure in the post is the same Provider Adapter condition at max reasoning. So 98.6% and 62.7% are the same model weights, same benchmark, two interfaces. ARC Prize said it will label the two conditions separately going forward. The New Stack reported that the harness producing the headline number is not itself part of the shipped product. Note that OpenAI's own launch page headlines 99.9%, not 98.6%, because its tables report the maximum score at any reasoning effort. ExploitBench. The 100% is confirmed in OpenAI's own preparedness document. That same document states OpenAI had contamination concerns about the benchmark, and therefore built a contamination-controlled internal port using high-severity V8 vulnerabilities disclosed between June and August 2026. On that controlled version, reporting of OpenAI's figures puts Astra at roughly 39.0% arbitrary code execution. OpenAI's own footnote further warns that some included vulnerabilities may not permit arbitrary code execution under the evaluation's constraints, so a 100% success rate may not be achievable in principle. FrontierMath Tier 4 (v2). The 97.6% is OpenAI's reported figure. Epoch AI's own benchmark page carries a standing disclosure that FrontierMath was developed with OpenAI funding and that OpenAI has exclusive access to a subset of the benchmark. Epoch also notes a June 2026 update that corrected errors in 42% of problems. No independent re-run of Astra on this benchmark was found. DeepSWE v1.1. This is the best-corroborated number in the claim. The independent Datacurve leaderboard lists Astra at roughly 74% with a stated uncertainty of ±3 in the shared mini-swe-agent setup, consistent with OpenAI's 74.1%. But the benchmark is 113 tasks, so one task is about 0.9 points, and the surrounding field is bunched: Claude Opus 5 at 73.7% and Gemini 3.8 Flash at 73.8%. Astra's margin over Opus 5 is 0.4 points, less than half a single task and far inside the published uncertainty band. Meta's Muse Spark 1.3 sits at 75.4% and was the leaderboard's number one on the claim's own publication date. On the comparative claim. OpenAI's own comparison table contains rows Astra loses. On Humanity's Last Exam with tools, Astra scores 57.2% against Claude Fable 5.1's 65.0%. On FrontierCode 1.1 splits, Fable 5 and Opus 5 edge Astra. Independently, Artificial Analysis measured Astra at about 61 on its Intelligence Index v4.1.1, level with GPT-5.6 Sol at 61 and behind Claude Fable 5.1 at roughly 66 and Opus 5 at roughly 63, while priced at 2.5 times Sol's rates. Artificial Analysis did find genuine gains for Astra in coding-agent performance and token efficiency. One contrary source deserves explicit handling. CryptoBriefing published a piece arguing the 98.6% claim was unverified and disconnected from a leaderboard topping out near 30%. The benchmark operator subsequently ran and published verified results confirming the figure under the stated harness. On this point the operator's own artifact governs, and that commentary is superseded.

Complete reasoning
As of 2026-09-04, the launch is real, and all four cited numbers exist in real, retrievable artifacts, so "False" and "Unverified" are both wrong, and refuting correct measurements would be an error rather than rigor. But the accurate family is closed by the contradiction test: the claim's operative proposition is that these scores show Astra outperforming GPT-5.6 Sol and Claude Fable 5.1, and OpenAI's own comparison table shows Fable 5.1 ahead on Humanity's Last Exam while the independent Artificial Analysis index puts Astra level with Sol and behind Fable 5.1 overall. That is a source I cite denying the operative proposition, not a simplification, so "Mostly accurate" is unavailable. I chose "Source exists but framing is misleading" over "Partially accurate but misleading" because the numbers are transcribed correctly rather than garbled: the entire distortion lives in stripped conditions, the harness for ARC-AGI-3, the contamination caveat for ExploitBench, the comparator choice and noise band for DeepSWE, and in an unscoped superiority verb. Confidence is High because the deciding artifacts were retrieved from the benchmark operators themselves, ARC Prize and Datacurve and Epoch AI, and from OpenAI's own documents, rather than resting on vendor claims alone.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/ddedfa1ec1ea/UwPS4ORsi-SgW8d4GJPpb3WhLP5

Ask this case

Answers come only from the case file above; nothing is added.

Did OpenAI actually release a model called GPT-6 Astra?

Yes. The launch on September 3, 2026 is confirmed by OpenAI's own launch page, system card, preparedness write-up and API documentation, plus a Microsoft Foundry availability post.

Are the four benchmark scores in the claim made up?

No, all four numbers are real and traceable to OpenAI's published materials, and the DeepSWE figure is independently corroborated. The issue is that each score carries a condition or caveat that was left out of the claim.

Why is the 98.6% ARC-AGI-3 score considered misleading?

That score came from a special OpenAI-supplied harness, while the benchmark's own operator, ARC Prize, ran the same model through its neutral setup and got 62.7%. Reporting indicates the harness behind the high score is not part of the product customers can actually buy.

What about the 100% score on ExploitBench?

OpenAI itself flagged that benchmark as possibly contaminated and built a corrected version to test for that. On the corrected version, the model scored around 39%, and OpenAI's own footnote says a 100% score may not even be achievable in principle under the evaluation's rules.

Does GPT-6 Astra really outperform Claude Fable 5.1?

The evidence contradicts this. OpenAI's own comparison table shows Fable 5.1 winning on Humanity's Last Exam, and independent benchmarking firm Artificial Analysis rated Astra level with its predecessor and behind Fable 5.1, while costing 2.5 times more.

Similar cases on record