Case TS-FA6D42E422 Aug 2026benchmark

AI

“NVIDIA's AVO agent system achieved a perfect 100.00 RHAE score on the public ARC-AGI-3 benchmark using Claude Opus 5, solving all 183 levels across 25 environments without any instruction, according to NVIDIA." Secondary claim on the post's image slide, investigated below: "NVIDIA's AI system scored a perfect 100 on ARC-AGI-3, solving…”

Plain restatementNVIDIA reports that its AVO agent system, running Anthropic's Claude Opus 5 as the underlying model, recorded a 100.00 RHAE score on the 25-environment public set of ARC-AGI-3, completing all 183 levels, with the agent given no stated rules or goals.

Mostly accurateConfidence Medium
What this verdict means →

Distortion codes this site does not recognise yet: benchmark_cherry_picking, harness_mismatch, cost_compute_omission. Not collectible until the field guide has an entry.

NVIDIA really did publish this. On August 21, 2026 its developer blog reported that its AVO agent system, running Anthropic's Claude Opus 5, scored 100.00 on the RHAE metric across all 183 levels of the 25 public ARC-AGI-3 environments. The Instagram caption's numbers are correct and it does say "public" and "according to NVIDIA," which is better than most coverage. The important missing context is that ARC Prize, which runs the benchmark, did not administer or verify this run, and that ARC Prize has said it will never report public-set scores on its official leaderboard because a harness built with knowledge of those environments can reach 100%. To prove that point ARC Prize itself released an open-source harness that scores 100% on the same public set by replaying human moves. NVIDIA states plainly that these are not results on the semi-private or private competition sets and that the jump from Claude Opus 5's separately measured 30.2% is not a controlled comparison. What remains unknown is whether AVO would perform anywhere near this on the held-out sets, and what the run cost in compute and time.

The drift / as claimed vs as evidenced

[drifted from the evidence:] NVIDIA's AVO agent system [drifted from the evidence:] achieved a [drifted from the evidence:] perfect 100.00 RHAE score on the public ARC-AGI-3 [drifted from the evidence:] benchmark using Claude Opus 5, solving all 183 levels [drifted from the evidence:] across 25 environments without any instruction, according to NVIDIA." Secondary claim on the [drifted from the evidence:] post's image slide, investigated below: "NVIDIA's AI system scored a perfect 100 on ARC-AGI-3, solving all 183 problems without any instruction.


[added by the neutral restatement:] NVIDIA reports that its AVO agent system, [added by the neutral restatement:] running Anthropic's Claude Opus 5 as the underlying model, recorded a 100.00 RHAE score on the [added by the neutral restatement:] 25-environment public [added by the neutral restatement:] set of ARC-AGI-3, [added by the neutral restatement:] completing all 183 levels, [added by the neutral restatement:] with the [added by the neutral restatement:] agent given no stated rules or goals.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
⌿ Omitted qualifier
A load-bearing condition from the source quietly disappears from the claim.
benchmark_cherry_picking
harness_mismatch
cost_compute_omission
Tertiary sourcecommunity commentary
Hacker News discussion thread
Secondary sourcetrade press
OfficeChai, "NVIDIA's Coding Agent AVO Scores 100% On ARC-AGI Benchmark"
Secondary sourcetrade press
The New Stack, "Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100%."
Primary sourcevendor
NVIDIA Technical Blog, "NVIDIA AVO Reaches 100% on ARC-AGI-3..." (vendor artifact of record for this result; retrieved via search index with verbatim passages)
Primary sourcebenchmark operator
ARC Prize, ARC-AGI-3 Technical Report (benchmark operator's rules on public-set scoring and task-specific overfitting)
Primary sourcebenchmark operator
ARC Prize results page, Claude Opus 5, 30.2% on ARC-AGI-3 as of July 24, 2026
Primary sourcebenchmark operator
ARC Prize Community Leaderboard repository and policy pages (self-reported scores, not verified, scorecard_url required for ARC-AGI-3)
Primary sourcevendor
NVIDIA AI official account post, August 21, 2026
● Primary source found
What is true
  • NVIDIA published this result on its own developer blog and its official AI account on August 21, 2026. The announcement is real, not fabricated.
  • The specific figures check out against NVIDIA's writeup: 100.00 RHAE, 183 levels, 25 environments, Claude Opus 5 as the underlying model.
  • The claim correctly says "public" and correctly attributes the result to NVIDIA. Both qualifiers matter and both are present.
  • The agent genuinely received no stated rules or goals within the episode, per NVIDIA's description of the setup.
  • The caption's "Claude Opus 5 alone was previously reported at roughly 30% under a different setup" is accurate, and "under a different setup" preserves the caveat NVIDIA itself makes.
What is misleading
  • Omitted qualifier: the claim reports a self-reported, unverified number without saying so. ARC Prize did not run or verify this result, and its own policy is that self-reported scores are untrustworthy by design. A reader hears "achieved" where the evidence supports "NVIDIA reports it achieved."
  • Benchmark cherry picking: the claim, and the coverage generally, presents a public-set result as a result on ARC-AGI-3. The public set is the one split the benchmark operator says it will never report on the official leaderboard, precisely because a harness built with knowledge of those environments can reach 100%. ARC Prize proved the point by publishing a human-replay harness that scores 100% on the same set. Without that context, a ceiling score on the practice set reads as a solved benchmark.
  • Harness mismatch: the implied 30% to 100% jump compares ARC Prize's administered run of Opus 5 through the official model interface against NVIDIA's agent running through NVIDIA's own reimplemented interface. NVIDIA states this is not a controlled ablation. The Instagram caption partially discloses this with "under a different setup," but the framing still invites the reader to attribute the entire gap to the AVO architecture.
  • Omitted qualifier, on the image slide specifically: the headline reads "scored a perfect 100 on ARC-AGI-3" with no mention of the public set, and says "183 problems" rather than 183 levels. On its own, that slide would earn a "source exists but framing is misleading" verdict, since it removes the single qualifier that determines what the number means.
  • Cost compute omission: neither the claim nor most coverage carries the resource picture. Environment action counts are published, but the token spend, dollar cost per environment, and wall-clock time for a long-horizon agent loop with persistent memory and supervision were not published in a form I could locate. ARC Prize's own leaderboards are built around cost per task for exactly this reason.
What is uncertain
  • Whether NVIDIA submitted a scorecard URL to the ARC-AGI Community Leaderboard, and whether ARC Prize has responded to or commented on the AVO result. No ARC Prize statement on it was found as of 2026-08-22.
  • Whether AVO's task interface or configuration was informed by knowledge of the public environments. NVIDIA says it reimplemented the task interface independently and adopted direct-interaction design principles from VISTA, but the writeup does not settle whether any tuning used public-environment knowledge, which is the exact condition ARC Prize defines as task-specific overfitting.
  • Whether the result would hold on the semi-private or private sets. Unknown, and NVIDIA does not claim it would.
  • Whether AVO's code for this run is public and reproducible. NVIDIA describes the architecture, but I could not confirm a release that would let a third party rerun it.
  • Full harness settings: reasoning effort, retry policy, number of attempts per environment, and cost.
Evidence summary

The underlying result is real and the numbers in the claim match NVIDIA's own writeup. NVIDIA states that using Claude Opus 5, AVO completed the full 25-environment public set with a 100.00 RHAE score, solving all 183 levels in 6,624 environment actions, and that VISTA reports 7,542 environment actions with Claude Opus 5 while completing the same 183 public-set levels, approximately 12% more. NVIDIA also states that this cross-system comparison should not be interpreted as a controlled ablation, because the two systems differ in agent backend, observation representation, memory, context management, and other implementation details. NVIDIA scopes the result itself. Its blog says the results cover the 25-environment ARC-AGI-3 public set using the official scorecard and RHAE metric, and are not results on the semi-private or fully private competition sets. The benchmark operator treats public-set scores as non-probative. The ARC-AGI-3 technical report defines task-specific overfitting to include any agent created with knowledge of public ARC-AGI-3 environments and then evaluated on those same environments, whether trained on them or using a harness handcrafted or specifically configured by someone with knowledge of the public environments; it states such agents can in principle achieve a 100% score on the public set, and that to demonstrate this ARC Prize released an open-source harness scoring 100% on all public environments using human replay; because it is impossible to ensure designers do not use the public environments in their work, and because the public set is materially easier than the private set, ARC Prize says it will never report public set scores of any system on the official leaderboard. The 30% comparator is real and is also a public-set number, but produced under ARC Prize's own administration. ARC Prize reports that as of July 24, 2026, Claude Opus 5 (High) is the highest-performing model on ARC-AGI-3 at 30.2%, having completed five additional Public Demo environments that no model had previously beaten. Trade coverage draws the distinction explicitly: the AVO figure is not a score on the ARC Prize leaderboard, since verified numbers there including Claude Opus 5's 30.2% come from runs the foundation independently administers, whereas AVO's result was generated and reported by NVIDIA's own team using its own reimplementation of the task interface. No independent verification of the AVO number was found. ARC Prize states that self-reported community scores are untrustworthy by design, that it does not run community leaderboard code or verify scores against the semi-private sets, and that for ARC-AGI-3 it collects a required scorecard_url and derives the score from it. Reaction was split: on X, many congratulated NVIDIA while others dismissed the result as overfitting on public data.

Complete reasoning
As of 2026-08-22, every number in the claim matches NVIDIA's own published writeup, and the claim retains the two qualifiers that most retellings drop: that this is the public set and that the source is NVIDIA. I considered "Source exists but framing is misleading" and rejected it for the caption-level claim, because the caption does not strip the decisive qualifier and even flags that the 30% comparator came from a different setup; that verdict would, however, be correct for the post's image headline, which drops "public" entirely. I considered "Accurate" and rejected it, because the claim omits that the score is unverified by ARC Prize and omits that the benchmark operator has published a human-replay harness scoring 100% on the same public set, context that changes how much the word "perfect" is worth. Confidence is Medium rather than High because the only evidence for the result is the vendor's own report, no independent runner or ARC Prize verification exists, and I could not confirm whether a scorecard was submitted for community inspection.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/fa6d42e4de6d/BLy7i-U-bRGMdMSp5_tz5kBTgnf

Ask this case

Answers come only from the case file above; nothing is added.

Did NVIDIA really claim its AI got a perfect 100 on ARC-AGI-3?

Yes. NVIDIA's developer blog and official AI account published this on August 21, 2026, reporting a 100.00 RHAE score across all 183 levels of the 25-environment public set using Claude Opus 5 as the underlying model.

Was this result verified by the people who run the benchmark?

No. ARC Prize, which administers ARC-AGI-3, did not run or verify this score. ARC Prize says self-reported community scores are untrustworthy by design and it does not check them against the semi-private sets.

Why does it matter that this was the 'public' set?

ARC Prize says it will never report public-set scores on its official leaderboard because a harness built with knowledge of those environments can reach 100 percent. ARC Prize proved this itself by releasing a human-replay harness that also scores 100 percent on the same public set.

Does this mean AVO is far better than Claude Opus 5 alone, which scored around 30%?

That comparison is not controlled. The 30.2 percent figure came from ARC Prize's own administered run through the official model interface, while NVIDIA's 100.00 figure came from NVIDIA's own reimplemented task interface. NVIDIA itself states this should not be read as a controlled ablation.

How would AVO perform on the harder, private test sets?

The case file does not establish this. NVIDIA's results only cover the public set and explicitly state they are not results on the semi-private or private competition sets, and no compute cost or wall-clock time figures for the run were found.

Similar cases on record