AI
“97.6% on FrontierMath Tier 4" / "97.6 on Frontier Math, Tier 4, the hardest math benchmark that exists" (Rowan Cheung, TikTok, sponsored video, 2026-09-17)”
Plain restatementOpenAI's GPT-6 Astra achieved a score of 97.6% on the Tier 4 split of Epoch AI's FrontierMath benchmark.
Distortion codes this site does not recognise yet: capability_extrapolation, benchmark_cherry_picking, harness_mismatch. Not collectible until the field guide has an entry.
The number is real. OpenAI's GPT-6 Astra did score 97.6% on FrontierMath Tier 4, and this is not just a company talking point: Epoch AI, the organisation that actually built and runs the benchmark, tested the model itself, got 98%, and has formally declared the benchmark saturated. Astra also solved the last Tier 4 problem no AI had cracked. Three things the video leaves out. First, the score is on a rewritten version of the test, issued in June 2026 after Epoch found errors in 42% of the problems and removed several. Second, Epoch openly discloses that OpenAI funded FrontierMath and has access to most of the problems and answers, which is worth knowing before treating any score on it as clean. Third, calling Tier 4 "the hardest math benchmark that exists" is already out of date: Epoch's newer Erdős benchmark of unsolved problems saw Astra solve just 2 of 68, a 3% score, and Astra actually trails Anthropic's Claude on Humanity's Last Exam. The specific claim holds up. The "benchmarks are saturated" story around it is selective, and the video is a disclosed paid promotion.
[drifted from the evidence:] 97.6% on FrontierMath Tier 4" / "97.6 on [drifted from the evidence:] Frontier Math, Tier 4, [drifted from the evidence:] the hardest math benchmark [drifted from the evidence:] that exists" (Rowan Cheung, TikTok, sponsored video, 2026-09-17)
[added by the neutral restatement:] OpenAI's GPT-6 Astra achieved a score of 97.6% on [added by the neutral restatement:] the Tier 4 [added by the neutral restatement:] split of Epoch AI's FrontierMath benchmark.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- GPT-6 Astra exists. OpenAI released it on 2026-09-03, and it is rolling out via ChatGPT paid plans and the API, matching the video's rollout statement.
- The 97.6% FrontierMath Tier 4 figure is genuinely in OpenAI's announcement table. It is not invented and it is not a misread of another row.
- The figure is independently corroborated. Epoch AI, the benchmark's operator, tested the model itself and reports 98% and formal saturation.
- Astra did solve the last unsolved Tier 4 problem, per Epoch, and Epoch confirms that solution was not a shortcut.
- The prime numbers claim in the video also checks out at the artifact level. OpenAI states Astra "improved a term in a bound on these gaps that had remained unchanged for more than 80 years" and published the proofs.
- Omitted qualifier: the claim says "FrontierMath Tier 4" with no version. The score is on Tier 4 v2, a 43-problem set issued in June 2026 after Epoch corrected errors in 42% of the original problems, with 7 problems removed entirely. A viewer hears a stable yardstick. The yardstick was rebuilt three months before the score.
- Marketing as evidence: the 97.6% digits come from OpenAI's own announcement table in a disclosed sponsored video. The independent corroboration exists and is genuine, but the video presents the vendor's number as simply the fact, with no indication of who ran it. The saturation finding belongs to Epoch. The 97.6% belongs to OpenAI.
- Omitted qualifier: the evaluator is not a disinterested party. Epoch discloses that OpenAI funded FrontierMath and has access to problem statements and solutions outside a 50-question holdout. That disclosure is material to any Tier 4 score and appears nowhere in the video.
- Capability extrapolation: "the hardest math benchmark that exists" is contradicted by the same operator's own newer artifact. Epoch launched FrontierMath Erdős, 68 historically unsolved problems requiring Lean solutions, and Astra scored 3%, solving 2 of 68. Every other model scored zero. Tier 4 was the hardest. It is not any more.
- Benchmark cherry picking: the video says frontier intelligence benchmarks are "basically already saturated." Astra scores 57.2% on Humanity's Last Exam with tools, behind Claude Fable 5.1 at 65.0%, Fable 5 at 63.8% and Opus 5 at 63.6%. On the independent Artificial Analysis Intelligence Index it sits at 61.2, effectively tied with its own predecessor and behind Fable 5.1 at 65.7. The saturated benchmarks are the ones shown.
- Internal inconsistency on the companion number: the spoken transcript says 99.9% on ARC-AGI-3 while the caption of the same post says 98.6%. OpenAI's page says 99.9%. The Decoder notes that 99.9% was achieved "under its own test conditions," and other coverage reports that the figure depends on a stateful OpenAI harness while independent stateless runs land far lower. This is a harness mismatch on the secondary number, not on the FrontierMath claim.
- I did not retrieve OpenAI's Academic comparison table directly. The exact digits 97.6 reach me through multiple independent secondary reproductions of that table, which agree with each other. The primary page I did reach states 98% in prose.
- The Tier 4 harness protocol is not established in the sources I reached: attempts per problem, whether the score is a single run or a mean across runs, and the reasoning-effort setting used for the specific 97.6% entry are not published in what I found. 42 of 43 is 97.7%, so 97.6% may be an average across runs rather than a single pass, but I will not reconstruct a protocol I did not see.
- Whether any contamination affected the result. OpenAI's access to Tier 4 problem statements and solutions outside the holdout is disclosed fact. Whether it bears on this score is not something the public evidence resolves either way.
- The mathematical results, including the prime gap bound, are published as Lean proofs. Commentary flags that the large proof file has not yet had independent human semantic review. I could not verify the review status.
The number is real and it is close to correct. OpenAI's own launch page states in prose that Astra "saturates FrontierMath Tier 4 with a 98% score." The more precise 97.6% figure comes from the Academic comparison table in that same announcement, which multiple independent analysts reproduce consistently: Vellum, DataCamp, The Decoder and Yotta Labs all record FrontierMath Tier 4 v2 at 97.6% for Astra against 83.0% for GPT-5.6 Sol and 87.8% for Claude Fable 5.1. Critically, this is not a vendor-only number. Epoch AI, which builds and operates FrontierMath, ran the model itself with pre-release access from OpenAI and reports its own result. Epoch states that Tier 4 launched on 2025-07-11 with a top score of 5%, that less than 14 months later the top score is 98%, and that it now considers the benchmark saturated. Epoch separately reports that GPT-6 Astra solved the final previously-unsolved Tier 4 problem, authored by Jay Pantone, and notes that unlike many Tier 4 problems this one was not solved via an unintended shortcut. Epoch's 98% and OpenAI's 97.6% are the same result at different rounding: Tier 4 v2 contains 43 problems, and 42 of 43 is 97.7%. So the operative proposition, that Astra scored 97.6% on FrontierMath Tier 4, is supported by the vendor's table and independently corroborated by the benchmark operator's own run.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/3750102e152a/5daJ8DyCSs08UvyivHDaMvc0lQq
Ask this case
Answers come only from the case file above; nothing is added.
Did GPT-6 Astra really score 97.6% on FrontierMath Tier 4?
Yes. The figure appears in OpenAI's own announcement table and is reproduced consistently by multiple independent analysts. Epoch AI, the organization that built and runs the benchmark, independently tested the model and reported 98%, which is the same result at different rounding.
Is the 97.6% score on the same test that has been used all along?
No. The score is on Tier 4 v2, a 43-problem version issued in June 2026 after Epoch found errors in 42% of the original problems and removed several. The claim does not mention that the test was rewritten just months before the score was achieved.
Is Epoch AI a neutral judge of this score?
Not entirely. Epoch discloses that OpenAI funded FrontierMath and has access to most of the problem statements and solutions outside a small holdout set. This conflict of interest is disclosed by Epoch but is not mentioned in the video.
Is Tier 4 really 'the hardest math benchmark that exists'?
That claim is out of date. Epoch's newer FrontierMath Erdős benchmark, made of unsolved problems, saw Astra solve only 2 of 68 problems, a 3% score, while every other model scored zero. Tier 4 was the hardest benchmark at one point but is not any more.
Does Astra lead on all major AI benchmarks, as the saturation claim implies?
No. On Humanity's Last Exam, Astra scores 57.2% with tools, behind Claude Fable 5.1, Fable 5, and Opus 5. The video highlights only the benchmarks where Astra leads.