AI
“GPT-5.4 xHigh tops Scale's standardized SEAL public board at 59.1%, the highest score any model reaches when every model runs the same harness on the public Pro set.”
Plain restatementOn Scale AI's standardized SWE-bench Pro public-set leaderboard, where all models are run under a common scaffold, GPT-5.4 at xHigh reasoning effort holds the top position with a 59.1% resolve rate, and no model on that board scores higher.
This was true in June 2026 and is out of date now. Scale AI's standardized SWE-bench Pro public leaderboard does list GPT-5.4 at xHigh effort with 59.1%, so the number itself is real and comes from an independent runner rather than from OpenAI. But the same board now shows Meta's Muse Spark 1.1 at 61.5%, added after that model launched on July 9 2026, so 59.1% is not the highest standardized score any more. The board's error bars are about plus or minus 3.5 points and its own ranking ties the two models at first place, so neither should be described as a clear winner. The phrase "any model" is also broader than the board is: many current frontier models, including the newest Claude, Gemini and GPT-5.6 releases, have no entry on it at all. One further thing changed since the claim was written. In July 2026 OpenAI published an audit finding roughly 30% of this benchmark's public tasks were broken and withdrew its recommendation of the benchmark, which weakens any ranking claim built on it. One detail I could not check: some top entries on the board carry an asterisk whose meaning I was unable to read, and if it flags differently run submissions it would qualify the "same harness" premise.
[drifted from the evidence:] GPT-5.4 xHigh tops Scale's standardized [drifted from the evidence:] SEAL public board at [drifted from the evidence:] 59.1%, the [drifted from the evidence:] highest score any model reaches when every model [drifted from the evidence:] runs the same harness on [drifted from the evidence:] the public Pro set.
[added by the neutral restatement:] On Scale AI's standardized [added by the neutral restatement:] SWE-bench Pro public-set leaderboard, where all models are run under a common scaffold, GPT-5.4 at [added by the neutral restatement:] xHigh reasoning effort holds the [added by the neutral restatement:] top position with a 59.1% resolve rate, and no model on [added by the neutral restatement:] that board scores higher.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- The number is real and correctly attributed. GPT-5.4 (xHigh) is listed at 59.10 on Scale's standardized SWE-bench Pro public-set board.
- GPT-5.4 (xHigh) is still co-ranked first on that board under the board's own confidence-interval tie rule.
- The claim correctly distinguishes the standardized board from vendor-scaffold numbers, which is the distinction most citations of "SWE-bench Pro" get wrong.
- "Same harness for every model" is an accurate description of what the standardized board does for the models it contains.
- Calling it a SEAL board is loose but not wrong. SWE-bench Pro is a Scale AI evaluation product, hosted alongside the SEAL leaderboards; the artifact of record is titled the SWE-Bench Pro Public Dataset leaderboard.
- The claim was accurate as a description of the June 2026 board, where GPT-5.4 (xHigh) led at 59.1% ahead of Muse Spark at 55.0%.
- Date context mismatch: the claim presents a June 2026 board state as the present one. As of 2026-08-12 the board's highest public-set score is Muse Spark 1.1 at 61.50, added after Meta's July 9 2026 release. "The highest score any model reaches" is no longer the board's arithmetic.
- Omitted qualifier: "any model" means "any model Scale has run." The standardized board contains no entry for Claude Fable 5, Opus 4.8, Opus 5, GPT-5.5, the GPT-5.6 line, GLM-5.2 or Kimi K3. The claim's universal quantifier describes a partial and lagging roster, not the model population.
- The claim treats a rank as a clean win when the board itself does not. GPT-5.4's ±3.56 interval overlaps Muse Spark 1.1's, and the board's numbering ties them at rank 1. Reading 59.1 as a distinct summit reads more precision off the board than the board publishes.
- Temporal overreach: the underlying runs are labelled by the benchmark operator as initial and subject to change, and the claim converts that provisional snapshot into a settled fact about model capability.
- Not a defect in the claim itself, but a change in what it means: since the claim was written, the frontier lab whose model it flatters audited this benchmark, found roughly 30% of the public tasks broken, and withdrew its recommendation of it. A leader-of-the-board statement about this set now carries much less weight than it did in June.
- Several top entries, including both Muse Spark rows, gpt-5.4 (xHigh), claude-opus-4-6 (thinking) and gemini-3.1-pro (thinking), carry an asterisk that older entries such as claude-opus-4-5-20251101 do not. I read the board's ranking table through the search index's rendering of the page and could not retrieve the footnote legend, so I cannot say what the asterisk denotes. If it marks provider-submitted or differently scaffolded runs, it would qualify the claim's "every model runs the same harness" premise directly. I am not asserting that it does.
- Whether the Muse Spark 1.1 entry of 61.50 is a Scale-run figure or a submitted one. Meta's own launch materials also report 61.5 on SWE-Bench Pro, an exact coincidence I could not resolve either way.
- The scaffold's precise identity. Scale's page text says SWE-Agent, while third-party descriptions and some Scale Labs board labels reference mini-SWE-agent. Which variant produced the 59.10 figure is not pinned.
- Whether Scale has revised, re-run, or annotated the public board in response to OpenAI's broken-task findings.
- The exact date Muse Spark 1.1 was added to the board, as distinct from its July 9 2026 model release date.
The deciding artifact is Scale's own public-set board, and its current state does not match the claim. The live public-set ranking reads Muse Spark 1.1 (NEW) at 61.50±3.10, gpt-5.4 (xHigh) at 59.10±3.56, Muse Spark at 55.00±3.60, claude-opus-4-6 (thinking) at 51.90±3.61, gemini-3.1-pro (thinking) at 46.10±3.60, claude-opus-4-5-20251101 at 45.89±3.60, descending through gpt-5-2025-08-07 (High) at 41.78, qwen3-coder-480b-a35b at 38.70, and gemma-3-27b-it at 11.38. Scale's leaderboard index shows the same top of board for the public open-source-repository track: Muse Spark 1.1 (NEW) 61.50±3.10, gpt-5.4 (xHigh) 59.10±3.56, Muse Spark 55.00±3.60. So the 59.1% figure is real and correctly attributed. The superlative attached to it is not: a higher standardized score exists on the same board. The rank numbering on the board interleaves as 1, 1, 3, 3, 5, 5, 5, indicating the board assigns tied ranks where confidence intervals overlap, which means GPT-5.4 (xHigh) is co-ranked first rather than the sole leader, and the model with the higher point score is Muse Spark 1.1. The change is datable. Meta released Muse Spark 1.1 on July 9, 2026, alongside a public preview of the Meta Model API, after the claim's June 2026 vintage. The claim's own origin chain confirms the staleness. The Morph LLM page that carries the sentence dates its table to June: "Scores below are from the public set (731 tasks), Pass@1, as of June 2026. GPT-5.4 (xHigh) leads at 59.1%, 4.1 points ahead of Meta's Muse Spark (55.0%) and 7.2 ahead of the best Claude run (Opus 4.6 thinking, 51.9%)." The same publisher's sibling page already reflects the newer board: "On Scale's standardized SWE-bench Pro board the newest frontier models are not run yet, so the top entries there are Muse Spark 1.1 (61.50%) and gpt-5.4 (59.10%). Scale SEAL public set (standardized scaffolding, July 2026): Muse Spark 1.1 61.50%..." One publisher, two pages, two different leaders. The claim's "same harness" premise is broadly faithful to how the board works. Scale states it ran frontier models on Pro using the SWE-Agent scaffold, and third-party analysis describes the standardized board as the place where "Scale's standardized SEAL public board is the only place every model runs the same harness". Run conditions are published: results are described as initial runs subject to change pending official announcement, models are run with uncapped cost and a turn limit of 250, and the ± column is a 95% binomial confidence interval over 730 problems. Separately, the benchmark's standing has deteriorated since the claim was written. OpenAI said it audited the Scale AI-developed benchmark and found a series of issues, and retracted earlier support after the audit found almost 30% of its tests were "broken". OpenAI's own announcement states it found 30% of SWE-Bench Pro tasks to be broken and is retracting its previous recommendation that the research community use it as a leading coding eval, noting that some correct solutions fail because of hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria. The audit used model-based investigator agents alongside independent reviews from five experienced software engineers.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/ebed4595dd92/xOgyGRi5NwpYjWBozEmFovJbpIE
Ask this case
Answers come only from the case file above; nothing is added.
Is the 59.1% score for GPT-5.4 xHigh real?
Yes. Scale AI's standardized SWE-bench Pro public-set board does list GPT-5.4 at xHigh reasoning effort with a 59.1% resolve rate, and that number is correctly attributed to an independent runner rather than to OpenAI.
So why is the claim rated superseded instead of true?
Because the board has changed since the claim was written. Meta's Muse Spark 1.1 was added on July 9, 2026 with a score of 61.5%, so 59.1% is no longer the highest score on that board.
Does that mean Muse Spark 1.1 is now the clear leader?
Not exactly. The board's own confidence intervals for GPT-5.4 and Muse Spark 1.1 overlap, and the board assigns them a tied rank of 1, so neither model should be called a clear winner.
Does the claim's phrase 'the highest score any model reaches' cover all current models?
No. It only covers models Scale has actually run on the board. Several current frontier models, including newer Claude, Gemini, and GPT-5.6 releases, have no entry on the board at all.
Has anything happened that affects how much this benchmark should be trusted at all?
Yes. In July 2026 OpenAI audited SWE-bench Pro, found that roughly 30% of its public tasks were broken, and withdrew its recommendation of the benchmark as a leading coding evaluation.