Case TS-096500D226 Sept 2026releaseCompound claim

AI

“A Anthropic lançou o Claude Opus 5.5, que rende no nível do Fable 5.1 (modelo mais potente da Anthropic) na maioria das tarefas e custa 40% menos para rodar do que o Opus 5.”

Plain restatementAnthropic has released a model called Claude Opus 5.5. It performs at the level of Claude Fable 5.1 on most tasks, and it costs 40 percent less to run than Claude Opus 5.

Mostly accurateConfidence Medium
What this verdict means →

Distortion codes this site does not recognise yet: cost_compute_omission, harness_mismatch. Not collectible until the field guide has an entry.

This post is largely accurate. Anthropic really did release Claude Opus 5.5 on September 22, 2026, and it is fully available now, not a preview. The two main statements in the post are almost word for word what Anthropic itself published: that the model performs at the level of Claude Fable 5.1 on most work and costs 40 percent less to run than Opus 5. The post also deserves credit for naming Anthropic as the source of its benchmark chart and for pointing out two tests where a rival model scored higher. The important missing detail is the 40 percent figure. Anthropic ties that number to default settings on typical workloads, and the actual published price cut is 20 percent on input and output tokens with cache reads down 60 percent. One independent measurement found that if you run the model at its highest effort setting, the cost saving over the previous model largely disappears, and that same independent test found a much narrower coding lead over the competing model than the launch chart shows.

The drift / as claimed vs as evidenced

[drifted from the evidence:] A Anthropic [drifted from the evidence:] lançou o Claude Opus 5.5, [drifted from the evidence:] que rende no nível do Fable 5.1 [drifted from the evidence:] (modelo mais potente da Anthropic) na maioria das tarefas e custa 40% [drifted from the evidence:] menos para rodar do que o Opus 5.


Anthropic [added by the neutral restatement:] has released a model called Claude Opus 5.5. [added by the neutral restatement:] It performs at the level of Claude Fable 5.1 [added by the neutral restatement:] on most tasks, and it costs 40 [added by the neutral restatement:] percent less to run than Claude Opus 5.

Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.

The trace / claim to source

Where it appeared
⌿ Omitted qualifier
A load-bearing condition from the source quietly disappears from the claim.
cost_compute_omission
harness_mismatch
$ Marketing as evidence
Promotional material dressed up as independent proof.
▲ Exaggeration
A real finding gets inflated: stronger, bigger, faster, or more certain than the evidence supports.
Secondary sourceindependent benchmark runner, read through third parties
Artificial Analysis measurements, as reported by Kingy AI, Latent Space AINews, Emergent and Implicator
Secondary sourcenamed-outlet journalism
TechCrunch, "Anthropic releases Opus 5.5 with lower prices and Fable-level performance"
Secondary sourcenamed-outlet journalism
VentureBeat launch coverage
Secondary sourcespecialist commentary
Vellum and Orca Router benchmark breakdowns
Primary sourcevendor official channel
Anthropic, "Introducing Claude Opus 5.5", official announcement and benchmark table
Primary sourcevendor official channel
Anthropic, Claude Opus product and pricing page
Primary sourcevendor official docs
Claude Platform pricing documentation
Primary sourcecloud provider official channel
AWS, "Claude Opus 5.5 is now available on AWS"
Primary sourcevendor official channel
Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1"
● Primary source found
What is true
  • Anthropic did release Claude Opus 5.5, on September 22, 2026, as the first model in a new Claude 5.5 family. It is generally available, not a preview or a waitlist.
  • Anthropic's own announcement says the model performs at the level of Claude Fable 5.1 on most work, which is what the post says.
  • Anthropic's own announcement says the model costs 40 percent less to run than Opus 5, which is what the post says.
  • The benchmark figures reproduced in the post match Anthropic's published launch table, and the post names Anthropic's table as the source rather than presenting the numbers as independent.
  • The post is correct that the model did not lead everywhere. On Anthropic's own table, GPT-6 Astra leads on AutomationBench at 41.4 percent against 40.0 percent, and on Terminal-Bench-Science at 64.6 percent against 58.7 percent.
  • An independent evaluator's aggregate index supports the performance-parity proposition, placing Opus 5.5 above Fable 5.1 on that index.
What is misleading
  • Omitted qualifier: the post states flatly that the model costs 40 percent less to run, and one slide says it simply "got 40 percent cheaper". Anthropic's own wording scopes the figure to default settings on typical workloads. The list price cut is 20 percent on input and output tokens, with cache reads down 60 percent, and the 40 percent figure blends those cuts with the model's token efficiency at its default effort level. A reader who takes it as a flat 40 percent price reduction has the wrong number.
  • Cost compute omission: the saving depends on the effort setting, and the post does not mention effort settings at all. Opus 5.5 defaults to medium effort while Opus 5 defaulted to high. At maximum effort, the independent per-task cost measurement reported for Opus 5.5 is roughly the same as Opus 5, so the advertised saving does not hold across the range of ways someone might actually run the model.
  • Harness mismatch: the slide showing Opus 5.5 leading GPT-6 Astra on agentic coding rests on a row where the two numbers were produced at different effort settings, with Astra's figure taken from OpenAI's own reporting. An independent runner using one harness measured the two as level on that benchmark. The lead is directionally real on the index level, but the size shown is a product of the comparison setup.
  • Marketing as evidence: although the post credits Anthropic's table, it presents vendor-run results as settling where the model leads, including a slide asserting leadership across every benchmark listed. Vendor numbers establish what the vendor claims, not what is independently true, and the one independent check available narrows several of those leads.
  • Omitted qualifier: the computer-use figure of 81.8 percent on OSWorld 2.0 is a partial-credit score. The strict completion score reported in the system card is 48.7 percent. Presenting only the partial figure under the heading of autonomous computer use overstates how often tasks are finished end to end.
What is uncertain
  • Whether Fable 5.1 is correctly described as Anthropic's most powerful model is genuinely ambiguous. Fable sits in a Mythos class that Anthropic positions above the Opus class, and Fable 5.1 was the prior public flagship. But in the same announcement the post draws from, Anthropic calls Opus 5.5 its new leading model, and there is also a restricted-access Mythos 5.1. This is a positioning and definition question rather than a factual error.
  • The Artificial Analysis figures, including the 59.6 percent Terminal-Bench result and the per-task cost decomposition, were read through third-party reports rather than from the evaluator's own pages, so the exact methodology behind them is not confirmed here.
  • Whether the 40 percent saving holds for any particular real workload is untestable from the published material, because "typical workloads" is not defined in the announcement.
  • No independent verification was found for the post's Opus 5 comparison figure on Humanity's Last Exam of 63.6 percent, although nothing contradicts it either.
Evidence summary

The release is real and the two headline propositions are near-verbatim translations of Anthropic's own published sentence. Anthropic's announcement page states that Opus 5.5 is the first model in the new Claude 5.5 family, that it "performs at the level of Claude Fable 5.1 on most work", and that it "costs 40% less to run than Opus 5". On cost, Anthropic's own page is more specific than the post. It states that "at default settings it will cost 40% less than Opus 5 on typical workloads". The list price change is smaller than 40 percent: input and output tokens are $4 and $20 per million, described by Anthropic as 20 percent below Opus 5, while cache reads fall to $0.20 per million, 60 percent below Opus 5. Anthropic also reports output generated more than 30 percent faster. The 40 percent figure is a blended estimate combining price cuts with token efficiency at default effort, not a list price cut. An independent evaluator, Artificial Analysis, was measured differently. As reported by several outlets, it placed Opus 5.5 at 58 on its Intelligence Index at maximum effort, ahead of both GPT-6 Astra and Fable 5.1 at 53. That corroborates the performance-parity proposition and arguably exceeds it. On cost, however, the same evaluator's decomposition reportedly found that at maximum effort Opus 5.5 cost about $5.98 per Intelligence Index task against about $5.86 for Opus 5 at maximum effort, meaning the per-task saving disappears at that setting. Opus 5.5 defaults to medium effort while Opus 5 defaulted to high, which is the comparison the 40 percent figure rests on. On the benchmark table the post reproduces, the figures match Anthropic's published numbers, and the post correctly names Anthropic as the source. The headline coding row carries an asymmetry: Anthropic's table reports Opus 5.5 at 66.4 percent on Terminal-Bench 4.0 run at extra-high effort against GPT-6 Astra at 57.9 percent taken at high effort from OpenAI's own figure. Artificial Analysis, running its own harness, reportedly measured Opus 5.5 at 59.6 percent on the same benchmark, level with Astra rather than ahead. Anthropic itself reports a standard error of plus or minus 2.6 points on that row, and on Terminal-Bench-Science the stated standard error is plus or minus 3.5 to 5.0 points, which is larger than the gap the post highlights on several other rows.

Complete reasoning
Both operative propositions, Fable-level performance on most tasks and 40 percent lower running cost than Opus 5, are near-verbatim renderings of Anthropic's own published sentence, the release is confirmed as generally available on the vendor's official channels as of 2026-09-25, and an independent evaluator's aggregate index corroborates the performance half rather than undercutting it. I considered "Accurate" and rejected it because the cost figure loses its scope in transmission: Anthropic ties it to default settings and typical workloads, and at maximum effort the independently reported per-task saving disappears. I considered "Partially accurate but misleading" and rejected it because the post credits Anthropic as the source of its numbers, discloses the two benchmarks where the model lost, and explicitly warns against hype, so no cited source contradicts the operative propositions and the framing does not reverse their meaning. Confidence is Medium rather than High because the comparative and cost claims rest mainly on vendor-run numbers, and the one independent check available diverges from the vendor table on the headline coding benchmark and on cost at non-default settings.
Use this case

The reply is formatted for pasting into the thread where the claim is circulating.

Compact share page: ai.trueseeker.com/s/096500d2b4f9/sAnSP95DPViXC3BtUspWYfl_vzA

Ask this case

Answers come only from the case file above; nothing is added.

Did Anthropic actually release a model called Claude Opus 5.5?

Yes. Anthropic released it on September 22, 2026, as the first model in a new Claude 5.5 family, and it is fully available now rather than a preview or waitlist product.

Is it true that Opus 5.5 costs 40 percent less to run than Opus 5?

Anthropic did publish that 40 percent figure, but it only applies at default settings on typical workloads. The actual list price cut is 20 percent on input and output tokens, and cache reads dropped 60 percent, with the 40 percent number blending those cuts with the model's default efficiency. An independent measurement found that at the highest effort setting, the cost saving largely disappears.

Does Opus 5.5 really perform at the level of Fable 5.1 on most tasks?

Anthropic's own announcement says so, and this matches what the post claims. An independent evaluator's aggregate index also placed Opus 5.5 above Fable 5.1, which supports this point.

Is Fable 5.1 really Anthropic's most powerful model?

This is unclear. Fable sits in a class Anthropic positions above Opus, but Anthropic's own announcement calls Opus 5.5 its new leading model, and there is also a more restricted Mythos 5.1. The case file treats this as a positioning question rather than a clear factual error.

Does Opus 5.5 beat competing models on every benchmark shown in the post?

No. On Anthropic's own table, GPT-6 Astra scores higher on AutomationBench and Terminal-Bench-Science. The post does correctly note that the model did not lead everywhere, though an independent test also found a narrower coding lead than the launch chart suggests.

Similar cases on record