AI
“Opus 5.5 might be the best coding model you can use right now. I let it run my Megabonk game test for almost 20 HOURS… and the result is kind of insane." Transcript: "I ran my Megabonk test on it and well, it worked for almost 20 hours. Building testing, building testing, building testing. And the result is pretty mind blowing. This is…”
Plain restatementA creator reports that Anthropic's Claude Opus 5.5 worked for close to 20 hours on his personal test of building a version of the game Megabonk, and that the output closely resembles the original game. Secondary claims in the same post: Opus 5.5 is available on Claude Pro, Max, Team and Enterprise plans starting at $20 per month, and costs $4 per million input tokens and $20 per million output tokens on the API.
Distortion codes this site does not recognise yet: cost_compute_omission, capability_extrapolation, benchmark_cherry_picking. Not collectible until the field guide has an entry.
The model in this post is real. Anthropic released Claude Opus 5.5 on September 22, 2026, and the prices quoted in the caption, $4 per million input tokens and $20 per million output tokens, match Anthropic's published rates, as does availability on the Pro, Max, Team and Enterprise plans. Megabonk is also a real game, a 3D roguelike released on Steam in September 2025. What cannot be checked is the test itself. The claim that the model worked for almost 20 hours and produced something close to the original rests only on the creator's own video, with no published prompt, code, playable build, session log or cost figure, and no one outside has inspected or reproduced it. Two details are missing that change the meaning: the post gives no token or dollar cost for a 20 hour run while calling it affordable at $20 per month, and it does not say whether original game art and sound were supplied to the model. Treat it as an unverified personal demonstration rather than an established result.
Opus 5.5 [drifted from the evidence:] might be the best coding model you can use right now. I let it run my Megabonk game test for [drifted from the evidence:] almost 20 HOURS… and the [drifted from the evidence:] result is [drifted from the evidence:] kind of insane." Transcript: "I ran my Megabonk test on [drifted from the evidence:] it and [drifted from the evidence:] well, it worked for almost 20 [drifted from the evidence:] hours. Building testing, building testing, building testing. And [drifted from the evidence:] the result is pretty mind blowing. This is the version it made and [drifted from the evidence:] it is pretty close to the [drifted from the evidence:] actual Megabonk game.
[added by the neutral restatement:] A creator reports that Anthropic's Claude Opus 5.5 [added by the neutral restatement:] worked for [added by the neutral restatement:] close to 20 hours [added by the neutral restatement:] on his personal test of building a version of the game Megabonk, and [added by the neutral restatement:] that the [added by the neutral restatement:] output closely resembles the original game. Secondary claims in the same post: Opus 5.5 is [added by the neutral restatement:] available on [added by the neutral restatement:] Claude Pro, Max, Team and [added by the neutral restatement:] Enterprise plans starting at $20 [added by the neutral restatement:] per month, and [added by the neutral restatement:] costs $4 per million input tokens and [added by the neutral restatement:] $20 per million output tokens on the [added by the neutral restatement:] API.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- Claude Opus 5.5 exists and was released by Anthropic on September 22, 2026, and the vendor positions it specifically for long-running agentic coding work.
- The API prices quoted in the caption, $4 per million input tokens and $20 per million output tokens, match Anthropic's published pricing as reported at launch.
- Opus 5.5 access on Pro, Max, Team and Enterprise plans, with Pro at $20 per month, matches launch-day reporting of Anthropic's plan availability.
- Megabonk is a real game, a 3D roguelike released on Steam on September 18, 2025 by the solo developer vedinad.
- Multi-hour coding runs by this model are a documented activity class. Anthropic's own page describes an Opus 5.5 code rewrite finishing in 9.5 hours, and other creators have published multi-hour Opus 5.5 game builds. These are vendor and creator self-reports, not independent evaluations.
- The post pairs a roughly 20 hour Opus 5.5 run with the words "surprisingly affordable" and the $20 per month Pro plan, but gives no token count and no dollar cost for the run itself. Claude Code enforces a rolling 5-hour window plus a weekly cap, so a 20 hour Opus workload is not a single straightforward Pro-plan session, and a reader cannot tell from the post what the test actually cost or on which access path it was run.
- "pretty close to the actual Megabonk game" is supported in the video by selected footage of a character roster, tier screens, sound and a pause menu. Visual and menu resemblance is not the same as matching a commercial game's play feel, balance, performance or content volume, and the post does not distinguish them.
- The setup is not described. Whether the model worked unattended, how many sessions and restarts were involved, how much human prompting or curation occurred, and whether original game assets or press-kit material were available to it are all unstated, and each materially changes what the 20 hour figure means.
- Benchmark cherry picking is not alleged here, but note the related gap: the superlative framing rests on one personal test plus a hedge, not on any evaluation set, and no comparison run on the same test with any other current model is shown in the post.
- Whether the run was close to 20 hours of continuous autonomous work or accumulated wall-clock time across multiple sessions. No log or timestamped artifact has been published.
- The total token spend and dollar cost of the run, which is unknown.
- How closely the output actually resembles Megabonk in play. There is no public build, no side-by-side comparison and no third-party playtest, so the resemblance claim cannot be checked by anyone outside the creator's footage.
- Whether any original Megabonk art, audio or other assets were supplied to the model, which bears directly on how much of the visual and audio similarity the model produced itself.
- Whether Opus 5.5 is in fact the strongest coding model available today. As of 2026-10-04 aggregators disagree: one places GPT-5.6 Sol at 58.9 on the Artificial Analysis index ahead of Opus 5.5 at 57.6, while two others place Opus 5.5 first. I did not retrieve the live board itself, and the post hedges this with "might be."
Claude Opus 5.5 is a real, shipped model. Anthropic announced it on September 22, 2026 as the first model in a Claude 5.5 family, positioned for long-running agentic coding, and says it performs at the level of Claude Fable 5.1 on most work while costing about 40 percent less to run than Opus 5. Launch reporting records the API model ID claude-opus-5-5, a 1 million token context window, and prices of $4 per million input tokens and $20 per million output tokens, with availability on the Claude platform, Amazon Bedrock, Google Cloud and Microsoft's cloud. Anthropic's own page cites long coding runs, including a rewrite task that Opus 5.5 finished in 9.5 hours against 12 hours for Fable 5.1, and quotes a customer describing a large multi-repository task left to run. Megabonk is a real commercial game, a 3D roguelike survival title by solo developer vedinad released on Steam on September 18, 2025. On the specific test described in the post, the only evidence that exists is the creator's own video and the newsletter and aggregator copies that republish his caption verbatim. No repository, prompt, playable build, session log, token count or cost figure has been published for this run, and no third party has played, inspected or reproduced the output. Other creators have published multi-hour Opus 5.5 game-building experiments, including Unreal Engine and roguelike builds, which show the general activity is real but do not test this claim.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/71a1df1af5b9/qvLqsQmoLPvH88je86uMMKO7IGf
Ask this case
Answers come only from the case file above; nothing is added.
Is Claude Opus 5.5 a real model, and did it really run for almost 20 hours on this test?
Claude Opus 5.5 is a real model released by Anthropic on September 22, 2026, and it is positioned for long-running coding work. However, the almost 20 hour run described in the post is only documented by the creator's own video, with no published log, timestamped artifact or other evidence to confirm the exact duration.
Does the result actually look close to the real Megabonk game?
The video shows selected footage of a character roster, tier screens, sound and a pause menu, which does resemble Megabonk visually and in its menus. But there is no public build, side-by-side comparison or third-party playtest, so whether it matches the actual game's play feel, balance or content cannot be checked.
How much did this 20 hour test cost to run?
The investigation did not establish this. The post calls the result affordable and cites the $20 per month Pro plan and standard API prices, but it gives no token count or dollar figure for the run itself, and Claude Code's usage caps mean a 20 hour workload is not a simple single Pro-plan session.
Were any original Megabonk assets, like art or sound, given to the model to help it build this?
This is unknown. The post does not say whether original game art, audio or other materials were supplied to the model, which matters because it would change how much of the resemblance the model actually created on its own.
Is Opus 5.5 really the best coding model available right now, as the title suggests?
This is not established. Leaderboard aggregators disagree, with one ranking a different model ahead of Opus 5.5 and two others ranking Opus 5.5 first, and the post itself hedges the claim with the word 'might'.