AI
“On the PosterBench benchmark, a mid-tier AI model equipped with the learned AutoDesign harness scored 71.83, beating a frontier model without the harness that scored 69.55”
Plain restatementIn the AutoDesign paper's PosterBench evaluation, a lower-cost model running with the paper's optimized DesignHarness recorded a score of 71.83, which is higher than the 69.55 recorded by a frontier model running without that harness.
Distortion codes this site does not recognise yet: benchmark_cherry_picking, harness_mismatch, capability_extrapolation, cost_compute_omission. Not collectible until the field guide has an entry.
The paper is real and both numbers are real. AutoDesign is a genuine preprint on arXiv from Meituan and university collaborators, and it does report a cheaper model with its optimized harness at 71.83 against a frontier model without that harness at 69.55. Two things are left out. Both scores come from a 10-paper subset called PosterBench-mini, not the paper's 100-paper main benchmark, and a 2.28-point gap on ten items is too small to settle an ordering. More importantly, the same table shows the frontier model running with that same harness scoring 74.56, higher than both, so the paper's actual result is that a better harness and a better model add together rather than one replacing the other. The benchmark, the scoring rubric, and the harness were all built by the same team, and no independent group has run it. The underlying idea that harness engineering delivers real gains is supported by the paper's data, but the specific "cheap model beats frontier model" framing is produced by choosing which frontier configuration to compare against.
[drifted from the evidence:] On the PosterBench [drifted from the evidence:] benchmark, a [drifted from the evidence:] mid-tier AI model [drifted from the evidence:] equipped with the [drifted from the evidence:] learned AutoDesign harness scored 71.83, [drifted from the evidence:] beating a frontier model without [drifted from the evidence:] the harness that [drifted from the evidence:] scored 69.55
[added by the neutral restatement:] In the [added by the neutral restatement:] AutoDesign paper's PosterBench [added by the neutral restatement:] evaluation, a [added by the neutral restatement:] lower-cost model [added by the neutral restatement:] running with the [added by the neutral restatement:] paper's optimized DesignHarness recorded a score of 71.83, [added by the neutral restatement:] which is higher than the 69.55 recorded by a frontier model [added by the neutral restatement:] running without that [added by the neutral restatement:] harness.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- The paper exists, is on arXiv as 2608.13560, and matches the screenshot's title, author list, and affiliations
- 71.83 is a real reported score, and it is a model running with DesignHarness attached
- 69.55 is a real reported score, and it is a baseline running without DesignHarness
- The two configurations use the same coding agent, Claude Code, and the same evaluation subset, so the pairing is not an apples-to-oranges harness comparison
- 71.83 is greater than 69.55. The arithmetic in the claim is correct
- The unnamed frontier model is Claude 4.8, a fair description of frontier
- The cheaper model is Doubao Seed 2.1 Pro at a reported $2.75 per poster against Claude 4.8 at $7.63, so "mid-tier" is defensible on price
- Several of the caption's secondary numbers check out: the +5.01 to +19.56 gain range across seven configurations; +5.59 for Codex with GPT-5.5 and +5.01 for Claude Code with Claude 4.8 as the two smallest gains; +19.56 as the largest, for Claude Code with DeepSeek V4 Pro; the 88 percent of top score at 27 percent of cost figure; and the roughly 18-point swing from changing only the coding harness, which is Table 3(b)'s Kimi Code at 82.31 against Claude Code at 64.33 with GLM 5.2 held fixed
- The caption's disclosure that the optimization loop plateaued and needed human redirection is accurate and appears in the paper's own Figure 2 caption
- Benchmark cherry picking: the claim selects the frontier model's un-harnessed configuration as the comparator, when the same table it draws 71.83 from shows Claude 4.8 with the same harness at 74.56. The win is produced by the choice of baseline. The paper's finding is that harness and model quality stack, not that one substitutes for the other. A reasonable reader takes away "fix your harness instead of upgrading your model," and the paper's own adjacent row says upgrading the model on top of the harness gains a further 2.73 points.
- Omitted qualifier: the claim says "on the PosterBench benchmark." Both figures are from PosterBench-mini, the fixed 10-paper subset, not the headline 100-paper Main Track. The paper is explicit about the distinction and the claim erases it. Ten items is a small enough set that a 2.28-point margin cannot carry the weight the claim puts on it.
- Marketing as evidence: the harness, the benchmark, the scoring rubric, and every number are all products of the same team, and the harness was optimized against feedback from this task. The paper takes real precautions, a held-out set the optimizer never sees and a system-blind human study, and those deserve credit, but they are internal controls, not independent evaluation. No third party has run PosterBench.
- Capability extrapolation: the caption generalizes from a 10-paper poster-design task to a purchasing rule about AI products in general, "some of what you pay frontier prices for is capability your engineering should be providing." The caption's own honest-scope paragraph concedes the single task domain, which partly offsets this, but the headline claim travels without that caveat.
- Cost compute omission: the "88 percent of the top score at 27 percent of the cost" figure is, in the paper's wording, a normalized designer-only API cost proxy. It is not the full cost of running the system, and the LongCat pricing footnote notes cached context was free on a cache hit at evaluation time. Two smaller notes that do not rise to distortions. First, "without the harness" is loose: native Claude Code is itself a harness, and the paper's baseline means without DesignHarness specifically, not harness-free. Second, "mid-tier" is the poster's word. Seed 2.1 Pro was the second-highest scoring of the models in Table 3(c), so it is mid-tier by price rather than by rank within the paper's own comparison set.
- Whether the paper reports variance, error bars, or repeated runs for PosterBench-mini. I read the paper through arXiv HTML and mirror excerpts rather than the complete PDF, and found no such reporting, but I cannot rule out that it appears in an appendix I did not reach
- The identity and settings of the VLM judge behind the PosterBench Score, and how the seven rubric dimensions are weighted into the composite
- Whether these results replicate. There is no independent reproduction, no critique, and no external run of PosterBench. The code is public, so replication is possible, but as of 2026-08-23 none has been published
- Whether the paper itself ever states the 71.83 versus 69.55 comparison. I found no passage making it. It appears to be assembled by the poster from two different tables, which is legitimate analysis but is the poster's inference, not the authors' claim
Both numbers are real and both appear in the paper. They are not invented and they are not mispaired. The 71.83 figure comes from Table 3(c), the Model Track, which the paper describes as fixing AutoDesign and Claude Code so as to "separate model choice from harness variation." In that table Claude 4.8 scores 74.56, Doubao Seed 2.1 Pro scores 71.83, Kimi K2.7 scores 70.12, and GLM 5.2 scores 64.33. Seed 2.1 Pro at 71.83 is therefore a model running with DesignHarness attached. The 69.55 figure comes from the PosterBench-mini main track. The paper states that AutoDesign "reaches 81.46 with Codex, compared with 75.87 for the native Codex baseline, and 74.56 with Claude Code, compared with 69.55 for the corresponding standalone baseline." The standalone baseline paired with the 74.56 Claude Code number is native Claude Code running Claude 4.8 without DesignHarness. This identification is confirmed arithmetically: 74.56 minus 69.55 equals 5.01, and the paper and repository both state the seven-configuration gain range as "+5.01 to +19.56 points," with 5.01 being the smallest. So the two numbers sit on the same evaluation set (PosterBench-mini, the 10-paper subset), use the same coding agent (Claude Code), and differ in model and in whether DesignHarness is attached. The inequality 71.83 > 69.55 is correct. The context the claim omits is in the same table. Claude 4.8 with DesignHarness scores 74.56, above Seed 2.1 Pro's 71.83. On the paper's own numbers, harness gains and model gains stack, and the frontier model still leads once both are given the same harness.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/e72ca05d7d81/Z_DXZDO7rzPrYDGlwL4uRxz18Z1
Ask this case
Answers come only from the case file above; nothing is added.
Are the numbers 71.83 and 69.55 real, or made up?
Both numbers are real and appear in the AutoDesign paper. They are not invented or mismatched, but they come from different tables and are being compared in a way the paper itself never states.
So did a mid-tier model actually beat a frontier model?
Only when the frontier model is compared without its own harness attached. The same table shows the frontier model, Claude 4.8, scoring 74.56 when given the same DesignHarness, which is higher than the mid-tier model's 71.83.
What is PosterBench-mini and why does it matter here?
It is a 10-paper subset used for this specific comparison, not the paper's full 100-paper main benchmark. A 2.28-point gap on just ten items is a small margin to base a strong claim on.
Who built the benchmark and does that affect the result?
The same team built the harness, the benchmark, and the scoring rubric, and no independent group has run PosterBench. The paper does include some internal controls like a held-out set and a human study, but these are not independent verification.
What is the investigation not able to confirm?
The investigation could not confirm whether the paper reports error bars or repeated runs, how the scoring rubric is weighted, or whether the results have been independently replicated. It also found no passage where the authors themselves make the 71.83 versus 69.55 comparison.