AI
“BDH-CQ, a 150M-parameter model that reasons silently in latent space instead of writing chain-of-thought text, scored 29.5% on ARC-AGI-1 at $0.0007 per task, and runs 11x cheaper than GPT-5.6 Luna (Low) even after OpenAI's 80% price cut, for about a 5-point accuracy tradeoff (29.5% vs 34.2%)”
Plain restatementA 150M-parameter recurrent model called BDH-CQ, which performs iterative computation in latent space rather than emitting chain-of-thought tokens, is reported to have scored 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, approximately 11 times cheaper per task than the GPT-5.6 Luna low-reasoning-effort configuration after OpenAI's 80% price reduction, with Luna (Low) scoring 34.2%.
Distortion codes this site does not recognise yet: benchmark_cherry_picking, harness_mismatch, cost_compute_omission, capability_extrapolation, eval_contamination, demo_to_product_conflation. Not collectible until the field guide has an entry.
The paper behind this post is real and the numbers are quoted correctly. A 150-million-parameter model called BDH-CQ, from the AI lab Pathway, did report scoring 29.5% on the public ARC-AGI-1 test at a computed cost of $0.0007 per task, and Pathway did publish the 11x cost comparison against GPT-5.6 Luna. The problem is what the comparison leaves out. Luna (Low) is the weakest setting of OpenAI's cheapest model tier, and the ARC Prize Foundation's own verified testing shows that same Luna model scoring 90.7% on ARC-AGI-1 at max effort, so the real accuracy gap is closer to 61 points than 5. The two cost figures are also measured differently: Pathway calculated its own from its hardware time at an assumed GPU rate with no profit margin, while Luna's comes from retail API pricing. Pathway itself says BDH-CQ is specialized for this kind of visual puzzle while the models it is compared to are general-purpose, so it is not a substitute for them at any price. The score is self-reported on a public test set rather than verified by the benchmark operator, and the paper reportedly does not publish a contamination check, so how much of the result reflects genuine reasoning versus familiarity with similar training data remains open.
[drifted from the evidence:] BDH-CQ, a 150M-parameter model [drifted from the evidence:] that reasons silently in latent space [drifted from the evidence:] instead of writing chain-of-thought [drifted from the evidence:] text, scored 29.5% on ARC-AGI-1 at $0.0007 per task, [drifted from the evidence:] and runs 11x cheaper than GPT-5.6 Luna [drifted from the evidence:] (Low) even after OpenAI's 80% price [drifted from the evidence:] cut, for about a 5-point accuracy tradeoff (29.5% vs 34.2%)
A 150M-parameter [added by the neutral restatement:] recurrent model [added by the neutral restatement:] called BDH-CQ, which performs iterative computation in latent space [added by the neutral restatement:] rather than emitting chain-of-thought [added by the neutral restatement:] tokens, is reported to have scored 29.5% on ARC-AGI-1 at [added by the neutral restatement:] a cost of $0.0007 per task, [added by the neutral restatement:] approximately 11 times cheaper [added by the neutral restatement:] per task than [added by the neutral restatement:] the GPT-5.6 Luna [added by the neutral restatement:] low-reasoning-effort configuration after OpenAI's 80% price [added by the neutral restatement:] reduction, with Luna (Low) scoring 34.2%.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- The paper exists and is correctly attributed. arXiv:2608.09888, submitted 10 August 2026, authors Engdahl, Kosowski, Chorowski and others, from Pathway.
- 150M parameters, latent-space iterative reasoning without verbalized chain-of-thought: accurately described.
- 29.5% and $0.0007 per task are the paper's actual headline figures, quoted correctly.
- The 34.2% figure for GPT-5.6 Luna (Low) and the roughly 11x post-price-cut cost ratio are exactly what Pathway published. The arithmetic checks: $0.040 × 0.20 ÷ $0.0007 = 11.4x.
- The 80% OpenAI price cut on GPT-5.6 Luna, dated 30 July 2026, is confirmed independently by ARC Prize.
- The claim states the accuracy gap rather than hiding it. Most aggregator coverage kept the multiplier and dropped the gap; this claim does not.
- The paper does contain the controlled-intervention analysis the post mentions.
- Benchmark cherry picking: the claim compares against GPT-5.6 Luna (Low), the lowest-reasoning-effort setting of OpenAI's cheapest tier, and presents the resulting 4.7-point gap as the accuracy cost of switching. The benchmark operator's verified figure for the same model after the same price cut is 90.7% on ARC-AGI-1 at $0.07 per task. Against that configuration the gap is about 61 points, not 5. The comparator choice, not the result, produces the "near-frontier" impression.
- Harness mismatch: BDH-CQ's 29.5% is self-reported on the 400-task public evaluation set. ARC Prize's verified Luna figures are on the semi-private set. The claim presents the two numbers as though they came off one board under one protocol.
- Cost compute omission: the two cost figures are computed on different bases and are not directly comparable. BDH-CQ's $0.0007 is Pathway's own hardware time at an assumed $3 per H200-hour, carrying no serving margin, no availability overhead, and no failed-call retries. Luna's figure is retail API pricing, which includes OpenAI's margin. Pathway discloses this; the claim does not. The post's slide text goes further, describing the computed cost as "enabling direct cross-model comparison," which is the opposite of what the disclosure supports.
- Omitted qualifier: "scored 29.5%" drops pass@2, drops that this is the highest of three effort settings (Low gives 21%), and drops that the score is self-reported rather than ARC Prize verified.
- Capability extrapolation: the surrounding post frames this as "near-frontier reasoning" and a "1000x smaller" model buying most of a frontier model's ability. Pathway's own page says BDH-CQ "is specialized for this kind of visual reasoning" while the leaderboard comparators are general-purpose systems. A model that does ARC grids and nothing else is not substitutable for GPT-5.6 Luna at any price, which is the practical meaning "11x cheaper" conveys to a reader.
- Eval contamination, unresolved rather than demonstrated: the training mixture reportedly blends privately curated examples with public ARC-derived datasets, with no published deduplication or contamination analysis, and the evaluation is on the public split. The post's slide asserts flatly that "ARC-AGI-1 is designed to resist memorization and test genuine few-shot reasoning," which describes the benchmark's design intent and does not establish that this particular result is uncontaminated.
- I retrieved the arXiv abstract page and the PDF's contributions section, but did not read the full paper end to end. The training-mixture composition, the three-effort-setting breakdown, and the reproducibility limitations come from secondary and tertiary analyses of the paper, not from my own reading of the full text. Those specific details are reported at Medium confidence.
- The independence of the two named reproductions cannot be settled from public sources. Pathway's blog presents Richard Zhong as a reproducer, while syndicated versions of the same release describe him as a co-author. If the auditors are co-authors, this is not independent verification.
- Whether the 34.2% Luna (Low) figure is a public-set or semi-private-set number is not resolvable from the pages I reached, which makes the exact comparability of the two accuracy figures indeterminate.
- No ARC Prize verified entry for BDH-CQ was found, and Pathway's page invites external API validation rather than reporting it. I state this as not found rather than as confirmed absence.
- Secondary claim, not investigated: reporting that Pathway raised additional funding at a $500M valuation, bringing total seed funding to $30M. That is a business claim requiring separate sourcing analysis.
The paper is real and the numbers are real. The arXiv abstract states directly: "A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency." The paper describes the mechanism the post describes: inputs presented at inference time continuously update the model's recurrent memory, and the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning. The vendor's own release carries the comparison the claim repeats: BDH-CQ scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.0007 per task, and runs approximately 11 times as cheaply per task as GPT 5.6 Luna (Low), even after accounting for OpenAI's 80% price cut of 5.6 Luna on July 30th, with Luna scoring 34.2% against BDH-CQ's 29.5%. The 80% price cut is independently confirmed by the benchmark operator: on July 30, 2026, OpenAI announced an 80% price reduction for GPT-5.6 Luna. Two facts in the primary sources sharply reframe the comparison. First, the cost figures are not measured the same way. Pathway itself discloses: Pathway's cost is computed from measured hardware time, while comparison costs are those reported to the leaderboard and may reflect API pricing for generalist models. The underlying rate is an assumption, not a price paid: inference takes roughly 0.85 H200 GPU-seconds per task at an assumed $3 per H200-hour, which is what pulls the price so low. Second, "Luna (Low)" is the weakest configuration of OpenAI's cheapest tier, not GPT-5.6's reasoning performance. The benchmark operator's own verified figure for the same model after the same price cut is far higher: at max reasoning effort under the new pricing, Luna scores 90.7% on ARC-AGI-1 Semi-Private at $0.07/task and 59.6% on ARC-AGI-2 Semi-Private at $0.18/task. Pathway also states the scope limit the post drops. On its own research page: "BDH-CQ is specialized for this kind of visual reasoning. While the commercial models on the leaderboard are general-purpose systems, the result matters because meaningful reasoning performance was achieved with a small fraction of the inference compute used by leading reasoning systems." On verification status, ARC Prize draws a bright line: scores are self-reported unless noted otherwise, results on the ARC-AGI-1 and ARC-AGI-2 semi-private sets are run and verified by ARC Prize, and everything else is scored on a public set and self-reported. BDH-CQ's 29.5% is a public-set number. Pathway's own page invites, rather than reports, API-level external validation.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/470187328400/3cMF5VQ3jGRdsJmtN74MxPyi7YL
Ask this case
Answers come only from the case file above; nothing is added.
Is the paper about BDH-CQ real?
Yes. It is arXiv:2608.09888, submitted August 2026 by authors from Pathway, and the model description, parameter count, and headline figures are quoted correctly from it.
So is the 5-point accuracy gap accurate?
Only for the specific comparison chosen. That gap is against GPT-5.6 Luna set to its lowest reasoning effort, but the benchmark operator's verified score for Luna at max effort is 90.7% on ARC-AGI-1, making the real gap closer to 61 points.
Are the two cost figures measured the same way?
No. BDH-CQ's $0.0007 is Pathway's own estimate from hardware time at an assumed GPU rate with no profit margin included, while Luna's cost comes from retail API pricing, which does include a margin. Pathway discloses this difference itself.
Was the 29.5% score independently verified?
No. It is a self-reported result on the public ARC-AGI-1 test set, not run or verified by the benchmark operator, unlike the semi-private-set figures ARC Prize verifies directly.
Does this mean BDH-CQ could replace a model like GPT-5.6 Luna?
The investigation did not establish that. Pathway itself states BDH-CQ is specialized for this type of visual puzzle, while the models it is compared against are general-purpose systems.