AI
“NVIDIA's AVO agent system achieved a perfect 100.00 RHAE score on the public ARC-AGI-3 benchmark using Claude Opus 5, solving all 183 levels across 25 environments without any instruction, according to NVIDIA." Secondary claim on the post's image slide, investigated below: "NVIDIA's AI system scored a perfect 100 on ARC-AGI-3, solving…”
Plain restatementNVIDIA reports that its AVO agent system, running Anthropic's Claude Opus 5 as the underlying model, recorded a 100.00 RHAE score on the 25-environment public set of ARC-AGI-3, completing all 183 levels, with the agent given no stated rules or goals.
Distortion codes this site does not recognise yet: benchmark_cherry_picking, harness_mismatch, cost_compute_omission. Not collectible until the field guide has an entry.
NVIDIA really did publish this. On August 21, 2026 its developer blog reported that its AVO agent system, running Anthropic's Claude Opus 5, scored 100.00 on the RHAE metric across all 183 levels of the 25 public ARC-AGI-3 environments. The Instagram caption's numbers are correct and it does say "public" and "according to NVIDIA," which is better than most coverage. The important missing context is that ARC Prize, which runs the benchmark, did not administer or verify this run, and that ARC Prize has said it will never report public-set scores on its official leaderboard because a harness built with knowledge of those environments can reach 100%. To prove that point ARC Prize itself released an open-source harness that scores 100% on the same public set by replaying human moves. NVIDIA states plainly that these are not results on the semi-private or private competition sets and that the jump from Claude Opus 5's separately measured 30.2% is not a controlled comparison. What remains unknown is whether AVO would perform anywhere near this on the held-out sets, and what the run cost in compute and time.
[drifted from the evidence:] NVIDIA's AVO agent system [drifted from the evidence:] achieved a [drifted from the evidence:] perfect 100.00 RHAE score on the public ARC-AGI-3 [drifted from the evidence:] benchmark using Claude Opus 5, solving all 183 levels [drifted from the evidence:] across 25 environments without any instruction, according to NVIDIA." Secondary claim on the [drifted from the evidence:] post's image slide, investigated below: "NVIDIA's AI system scored a perfect 100 on ARC-AGI-3, solving all 183 problems without any instruction.
[added by the neutral restatement:] NVIDIA reports that its AVO agent system, [added by the neutral restatement:] running Anthropic's Claude Opus 5 as the underlying model, recorded a 100.00 RHAE score on the [added by the neutral restatement:] 25-environment public [added by the neutral restatement:] set of ARC-AGI-3, [added by the neutral restatement:] completing all 183 levels, [added by the neutral restatement:] with the [added by the neutral restatement:] agent given no stated rules or goals.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- NVIDIA published this result on its own developer blog and its official AI account on August 21, 2026. The announcement is real, not fabricated.
- The specific figures check out against NVIDIA's writeup: 100.00 RHAE, 183 levels, 25 environments, Claude Opus 5 as the underlying model.
- The claim correctly says "public" and correctly attributes the result to NVIDIA. Both qualifiers matter and both are present.
- The agent genuinely received no stated rules or goals within the episode, per NVIDIA's description of the setup.
- The caption's "Claude Opus 5 alone was previously reported at roughly 30% under a different setup" is accurate, and "under a different setup" preserves the caveat NVIDIA itself makes.
- Omitted qualifier: the claim reports a self-reported, unverified number without saying so. ARC Prize did not run or verify this result, and its own policy is that self-reported scores are untrustworthy by design. A reader hears "achieved" where the evidence supports "NVIDIA reports it achieved."
- Benchmark cherry picking: the claim, and the coverage generally, presents a public-set result as a result on ARC-AGI-3. The public set is the one split the benchmark operator says it will never report on the official leaderboard, precisely because a harness built with knowledge of those environments can reach 100%. ARC Prize proved the point by publishing a human-replay harness that scores 100% on the same set. Without that context, a ceiling score on the practice set reads as a solved benchmark.
- Harness mismatch: the implied 30% to 100% jump compares ARC Prize's administered run of Opus 5 through the official model interface against NVIDIA's agent running through NVIDIA's own reimplemented interface. NVIDIA states this is not a controlled ablation. The Instagram caption partially discloses this with "under a different setup," but the framing still invites the reader to attribute the entire gap to the AVO architecture.
- Omitted qualifier, on the image slide specifically: the headline reads "scored a perfect 100 on ARC-AGI-3" with no mention of the public set, and says "183 problems" rather than 183 levels. On its own, that slide would earn a "source exists but framing is misleading" verdict, since it removes the single qualifier that determines what the number means.
- Cost compute omission: neither the claim nor most coverage carries the resource picture. Environment action counts are published, but the token spend, dollar cost per environment, and wall-clock time for a long-horizon agent loop with persistent memory and supervision were not published in a form I could locate. ARC Prize's own leaderboards are built around cost per task for exactly this reason.
- Whether NVIDIA submitted a scorecard URL to the ARC-AGI Community Leaderboard, and whether ARC Prize has responded to or commented on the AVO result. No ARC Prize statement on it was found as of 2026-08-22.
- Whether AVO's task interface or configuration was informed by knowledge of the public environments. NVIDIA says it reimplemented the task interface independently and adopted direct-interaction design principles from VISTA, but the writeup does not settle whether any tuning used public-environment knowledge, which is the exact condition ARC Prize defines as task-specific overfitting.
- Whether the result would hold on the semi-private or private sets. Unknown, and NVIDIA does not claim it would.
- Whether AVO's code for this run is public and reproducible. NVIDIA describes the architecture, but I could not confirm a release that would let a third party rerun it.
- Full harness settings: reasoning effort, retry policy, number of attempts per environment, and cost.
The underlying result is real and the numbers in the claim match NVIDIA's own writeup. NVIDIA states that using Claude Opus 5, AVO completed the full 25-environment public set with a 100.00 RHAE score, solving all 183 levels in 6,624 environment actions, and that VISTA reports 7,542 environment actions with Claude Opus 5 while completing the same 183 public-set levels, approximately 12% more. NVIDIA also states that this cross-system comparison should not be interpreted as a controlled ablation, because the two systems differ in agent backend, observation representation, memory, context management, and other implementation details. NVIDIA scopes the result itself. Its blog says the results cover the 25-environment ARC-AGI-3 public set using the official scorecard and RHAE metric, and are not results on the semi-private or fully private competition sets. The benchmark operator treats public-set scores as non-probative. The ARC-AGI-3 technical report defines task-specific overfitting to include any agent created with knowledge of public ARC-AGI-3 environments and then evaluated on those same environments, whether trained on them or using a harness handcrafted or specifically configured by someone with knowledge of the public environments; it states such agents can in principle achieve a 100% score on the public set, and that to demonstrate this ARC Prize released an open-source harness scoring 100% on all public environments using human replay; because it is impossible to ensure designers do not use the public environments in their work, and because the public set is materially easier than the private set, ARC Prize says it will never report public set scores of any system on the official leaderboard. The 30% comparator is real and is also a public-set number, but produced under ARC Prize's own administration. ARC Prize reports that as of July 24, 2026, Claude Opus 5 (High) is the highest-performing model on ARC-AGI-3 at 30.2%, having completed five additional Public Demo environments that no model had previously beaten. Trade coverage draws the distinction explicitly: the AVO figure is not a score on the ARC Prize leaderboard, since verified numbers there including Claude Opus 5's 30.2% come from runs the foundation independently administers, whereas AVO's result was generated and reported by NVIDIA's own team using its own reimplementation of the task interface. No independent verification of the AVO number was found. ARC Prize states that self-reported community scores are untrustworthy by design, that it does not run community leaderboard code or verify scores against the semi-private sets, and that for ARC-AGI-3 it collects a required scorecard_url and derives the score from it. Reaction was split: on X, many congratulated NVIDIA while others dismissed the result as overfitting on public data.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/fa6d42e4de6d/BLy7i-U-bRGMdMSp5_tz5kBTgnf
Ask this case
Answers come only from the case file above; nothing is added.
Did NVIDIA really claim its AI got a perfect 100 on ARC-AGI-3?
Yes. NVIDIA's developer blog and official AI account published this on August 21, 2026, reporting a 100.00 RHAE score across all 183 levels of the 25-environment public set using Claude Opus 5 as the underlying model.
Was this result verified by the people who run the benchmark?
No. ARC Prize, which administers ARC-AGI-3, did not run or verify this score. ARC Prize says self-reported community scores are untrustworthy by design and it does not check them against the semi-private sets.
Why does it matter that this was the 'public' set?
ARC Prize says it will never report public-set scores on its official leaderboard because a harness built with knowledge of those environments can reach 100 percent. ARC Prize proved this itself by releasing a human-replay harness that also scores 100 percent on the same public set.
Does this mean AVO is far better than Claude Opus 5 alone, which scored around 30%?
That comparison is not controlled. The 30.2 percent figure came from ARC Prize's own administered run through the official model interface, while NVIDIA's 100.00 figure came from NVIDIA's own reimplemented task interface. NVIDIA itself states this should not be read as a controlled ablation.
How would AVO perform on the harder, private test sets?
The case file does not establish this. NVIDIA's results only cover the public set and explicitly state they are not results on the semi-private or private competition sets, and no compute cost or wall-clock time figures for the run were found.