AI
“Anthropic launched Claude Opus 5.5, the first model in its new Claude 5.5 family, which performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
Plain restatementAnthropic released a model called Claude Opus 5.5 on or around 2026-09-22, described as the first entry in a Claude 5.5 model family. Anthropic states its performance is comparable to Claude Fable 5.1 across most tasks and that running it costs approximately 40% less than Claude Opus 5.
Distortion codes this site does not recognise yet: harness_mismatch, benchmark_cherry_picking. Not collectible until the field guide has an entry.
Anthropic did launch Claude Opus 5.5 on September 22, 2026, and it is the first model in a new Claude 5.5 family. The model is genuinely available now through Anthropic's API, the Claude apps on paid plans, and Amazon, Google and Microsoft cloud platforms. The performance and cost sentence in the post is copied almost word for word from Anthropic's own announcement, so it is not something the poster invented, but it is the company's own claim about its own product rather than an independent finding. The price cut is real and documented at $4 and $20 per million input and output tokens, down 20%, with cache reads down 60%, and the headline 40% figure is Anthropic's estimate for typical workloads that also assumes the model uses fewer tokens per task. One detail worth knowing is that the benchmark scores and the cost saving are measured at different settings, so no single run delivers both at once. Independent testers broadly agree the model is at or above the Fable 5.1 level and is much cheaper to run, though at least two found areas where Fable 5.1 still edges it out. Anthropic itself notes the gap between the two models is narrower than its benchmark table suggests, and that caveat did not make it into the post.
Anthropic [drifted from the evidence:] launched Claude Opus 5.5, the first [drifted from the evidence:] model in [drifted from the evidence:] its new Claude 5.5 family, [drifted from the evidence:] which performs at the level of Claude Fable 5.1 [drifted from the evidence:] on most [drifted from the evidence:] work and costs 40% less [drifted from the evidence:] to run than Opus 5.
Anthropic [added by the neutral restatement:] released a model called Claude Opus 5.5 [added by the neutral restatement:] on or around 2026-09-22, described as the first [added by the neutral restatement:] entry in [added by the neutral restatement:] a Claude 5.5 [added by the neutral restatement:] model family. [added by the neutral restatement:] Anthropic states its performance is comparable to Claude Fable 5.1 [added by the neutral restatement:] across most [added by the neutral restatement:] tasks and [added by the neutral restatement:] that running it costs [added by the neutral restatement:] approximately 40% less than [added by the neutral restatement:] Claude Opus 5.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- Anthropic did launch Claude Opus 5.5 on 2026-09-22. This is confirmed on Anthropic's own launch page, product page, and platform documentation, and reflected in the Opus 5 docs' migration notice.
- It is the first model in a new Claude 5.5 family. Anthropic has signalled Sonnet 5.5 and Haiku 5.5 to follow.
- The model is genuinely shipped and generally available, not announced, waitlisted, or in preview. It is live in consumer apps on paid tiers and across three major clouds plus Anthropic's own API.
- The sentence quoted in the claim is Anthropic's own wording, reproduced essentially verbatim rather than invented or embellished by the post.
- Prices did fall. The 20% per-token cut and 60% cache-read cut are on the vendor's pricing page of record.
- The direction of the performance claim is independently corroborated at the top line. Artificial Analysis and Snorkel both place Opus 5.5 at or above the Fable 5.1 level on their own harnesses.
- The post's secondary claims are accurate. METR did publish an independent predeployment evaluation, and the "strongest-performing model on our automated behavioral audit" line is an accurate quotation of Anthropic's self-report.
- Marketing as evidence: the claim states "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run" as plain fact, in the investigator's intake wording, with no attribution. Both halves originate entirely with Anthropic. The post's own caption is better behaved than the extracted claim line, since it writes "the company says" and credits "Source: Anthropic," but the headline formulation drops that framing. Under the vendor duality rule, Anthropic's page is decisive evidence that Opus 5.5 exists, is available, and costs $4/$20. It is an interested party's self-report on whether it matches Fable 5.1.
- Omitted qualifier: Anthropic's own page says the 40% figure is an *estimated* saving for *typical workloads billed by token*. The claim drops "estimated" and "typical workloads." The saving is a modelled blend of a 20% rate cut, a 60% cache-read cut, and an assumption that the model uses fewer tokens per task. A workload with low cache-hit rates, or one that raises effort above the new medium default, will not see 40%.
- Harness mismatch: the headline scores and the headline saving are not measured under the same configuration. Most benchmark rows run Opus 5.5 at max effort, while the 40% saving is computed at each model's default effort, which changed from high on Opus 5 to medium on Opus 5.5. No single run produces both the quoted performance and the quoted cost at once. Separately, the GPT-6 Astra and GPT-5.6 Sol columns in the same table are OpenAI-reported numbers rather than figures Anthropic reproduced.
- Exaggeration by omission of the vendor's own hedge: Anthropic explicitly cautions that benchmark margins are now a weaker guide to real-world differences and that the Opus 5.5 versus Fable 5.1 gap is narrower than its table suggests. The post carries the favourable table across four slides and leaves the caveat behind. This is unusual in that the qualifier being stripped is the vendor's own.
- Benchmark cherry picking, mild: independent results are less uniform than the launch framing. Snorkel's pass@1 is effectively tied, with Fable 5.1 nominally ahead, and its pass@5 puts Opus 5.5 below Opus 5. Endor Labs ranks Opus 5.5 third on secure code generation, behind Fable 5.1, and raises a memorization question about the low tool-call count. None of this refutes "performs at the level of," which is a parity claim rather than a superiority claim, but the post's presentation reads as a clean sweep.
- Whether "performs at the level of Fable 5.1 on most work" holds for real user workloads rather than benchmark suites. This is unresolvable in the two days since launch, and Anthropic itself declines to claim it strongly.
- Whether the 40% saving reproduces outside Anthropic's modelled typical workload. Endor Labs' independent cost measurement is directionally supportive and one comparison shopper put the real-world gap against Fable 5.1 at 34% rather than the 60% list-price gap, but I found no independent replication of the specific 40%-versus-Opus-5 figure at default settings.
- The Endor Labs memorization finding is a single evaluator's observation and has not been replicated elsewhere as of 2026-09-24.
- I reviewed the Anthropic, Artificial Analysis, METR, Snorkel and Endor Labs pages through retrieved page content rather than by loading each in full. I did not independently verify the full Anthropic benchmark table row by row against the live page.
- The post's claim that Opus 5.5 "is now available across Claude products" is imprecise: Free-tier accounts do not have it. This is a minor scope overstatement that does not affect the primary claim.
The claim is not a paraphrase. It is a near-verbatim reproduction of Anthropic's own launch sentence. Anthropic's official page states: "It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." The company's own social announcement used slightly different wording, describing the model as performing at Fable 5.1's level "for most tasks." The release itself is verified on the vendor's own channels. Opus 5.5 is available on Claude for Pro, Max, Team, and Enterprise users, and available to developers on the Claude Platform natively and on Amazon Web Services, Google Cloud, and Microsoft Foundry. Free-tier accounts are excluded. Anthropic's own Opus 5 documentation now carries a migration notice: "Although Claude Opus 5 is still available, you should consider migrating to Claude Opus 5.5 for improved performance." Anthropic has said Sonnet 5.5 and Haiku 5.5 are expected to follow, which supports the "first model in the family" framing. On cost, the vendor page states: pricing for Opus 5.5 will cost an estimated 40% less to run than Opus 5 for typical workloads billed by token, with Opus 5.5 at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5, and cache reads 60% lower at $0.20 per million tokens. The 40% figure is therefore a composite estimate, not a list-price cut. Anthropic states that Opus 5.5 requires less compute to serve than Opus 5 and that its pricing reflects that. Independent framing of the same point: the model is described as running about 40% cheaper on typical workloads not just from the lower rate but because it finishes work in fewer tokens and fewer turns. On performance parity, Anthropic's published table shows Opus 5.5 ahead of Fable 5.1 on several agentic benchmarks. Anthropic reports 66.4% on Terminal-Bench 4.0 against 55.8% for Fable 5.1 and 52.3% for Opus 5, and 54.4% on FrontierCode v1.1 Main against 50.3% for Fable 5.1 and 48.0% for Opus 5. Anthropic also reports cross-effort comparisons: at max effort Opus 5.5 scores 1846 Elo where Fable 5.1 scores 1735 and Opus 5 scores 1708, and on CursorBench at default medium effort Opus 5.5 scores 52.5% compared to 51.8% for Fable 5.1 at max. Independent evaluation partly corroborates and partly complicates the parity claim. Artificial Analysis, running its own harness, records a score of 58 on its Intelligence Index for Opus 5.5 at max effort, with 54 at high effort and 51 at medium effort, described as well above the median of 25 for reasoning models in a similar price tier. Snorkel AI, which has run the same expert-built task set across three model generations, reports a pass rate of 68% for Opus 5.5 on its frontier task set, up from 61% for Opus 5 and 49% for Fable 5.1 on the same tasks, but also that pass@1 is 61.5% for Fable 5.1, 60.7% for Opus 5, and 60.7% for Opus 5.5, a flattening across the last two generations with Fable 5.1 nominally higher than the other two. Snorkel separately notes pass@5 of 74.1% for Fable 5.1, 79.3% for Opus 5 and 76.7% for Opus 5.5, where Opus 5.5 falls short of Opus 5. Endor Labs found Opus 5.5 ranking third on generating secure code, with Claude Code running Fable 5.1 at the top, while also finding it the most cost-efficient Anthropic model they have run, by a wide margin. Notably, Anthropic itself attaches a hedge that most vendors omit: per analyst writeups of the launch table, Anthropic adds that at this level of capability benchmark margins have become a less reliable guide to real-world differences, and that the gap between Opus 5.5 and Fable 5.1 is narrower than the scores suggest. The post's secondary claims also check out. METR published a summary of its independent predeployment evaluation of Claude Opus 5.5 on September 22, 2026, and METR states its work was oriented around collecting evidence related to AI R&D capabilities and was not meant to verify compliance with any specific threshold from Anthropic's policies, nor to assess alignment properties. The "strongest-performing model on our automated behavioral audit" line is accurately quoted from Anthropic's own page, and is a vendor self-report.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/5f21f22af3c5/cM0Mf3ERs8hbqKZNZ6rA5MGoLxo
Ask this case
Answers come only from the case file above; nothing is added.
Did Anthropic actually release Claude Opus 5.5?
Yes. Anthropic launched it on or around September 22, 2026, and it is genuinely available through Anthropic's API, the Claude apps on paid plans, and Amazon, Google and Microsoft cloud platforms, not just announced or waitlisted.
Is it true that Opus 5.5 performs at the level of Claude Fable 5.1 on most work?
This is Anthropic's own claim about its own product, taken almost word for word from the company's launch page. Independent testers like Artificial Analysis and Snorkel AI broadly place Opus 5.5 at or above the Fable 5.1 level on their own tests, though Snorkel and Endor Labs also found specific measures where Fable 5.1 still does better.
Does Opus 5.5 really cost 40% less to run than Opus 5?
The listed price drop is 20% per token and 60% for cache reads, both documented on Anthropic's pricing page. The 40% figure is Anthropic's own estimate for typical workloads, which also assumes the model uses fewer tokens per task, so it is not a straightforward price cut.
Are the performance and cost claims measured under the same conditions?
No. The case file notes that the benchmark scores and the cost saving are measured at different settings, so no single run demonstrates both the performance parity and the cost saving at once.
Did Anthropic mention any limits to the performance comparison?
Yes. Anthropic itself notes that at this level of capability, benchmark margins are a less reliable guide to real-world differences, and that the actual gap between Opus 5.5 and Fable 5.1 is narrower than its benchmark table suggests. That caveat was not included in the post being fact-checked.