AI
“GLM-5.3 uses the same base model as GLM-5.2, with performance gains coming almost entirely from post-training: Z.ai Code Bench improved from 20.9% to 31.4%, Terminal-Bench 3.0 from 4.6% to 28.3%, DeepSWE from 46.2% to 66.9%, and AutomationBench from 26.2% to 48.2%.”
Plain restatementZ.ai's GLM-5.3 is built on the unchanged GLM-5.2 base model, and Z.ai attributes its benchmark gains to expanded post-training. Z.ai reports four specific score pairs: Z.ai Code Bench 20.9 to 31.4, Terminal-Bench 3.0 4.6 to 28.3, DeepSWE 46.2 to 66.9, AutomationBench 26.2 to 48.2.
Distortion codes this site does not recognise yet: harness_mismatch, benchmark_cherry_picking. Not collectible until the field guide has an entry.
The numbers in this post are accurately copied from Z.ai's own GLM-5.3 launch materials, published 14 August 2026. Z.ai does say GLM-5.3 reuses the GLM-5.2 base model unchanged and that all gains come from expanded post-training, and Bloomberg reported the same. The four score pairs check out against Z.ai's documentation and launch table. The important missing context is that Z.ai ran all four of these evaluations itself, inside its own harness, and no independent party has reproduced any GLM-5.3 result, partly because the model weights are being held back for about two weeks over cybersecurity concerns. Two smaller caveats: the Z.ai Code Bench figures of 20.9 to 31.4 are the "High effort" setting on a chart with several settings, and that benchmark is private and cannot be rerun by anyone outside the company. The Terminal-Bench jump also starts from a near-zero baseline of 4.6, which makes the multiple look larger than the underlying progress. Z.ai's same table also shows GLM-5.3 losing to GPT-5.6 Sol and Claude Fable 5 on several of these benchmarks, which the post does not mention.
GLM-5.3 [drifted from the evidence:] uses the [drifted from the evidence:] same base model [drifted from the evidence:] as GLM-5.2, with performance gains [drifted from the evidence:] coming almost entirely from post-training: Z.ai Code Bench [drifted from the evidence:] improved from 20.9% to 31.4%, Terminal-Bench 3.0 [drifted from the evidence:] from 4.6% to 28.3%, DeepSWE [drifted from the evidence:] from 46.2% to 66.9%, [drifted from the evidence:] and AutomationBench [drifted from the evidence:] from 26.2% to 48.2%.
[added by the neutral restatement:] Z.ai's GLM-5.3 [added by the neutral restatement:] is built on the [added by the neutral restatement:] unchanged GLM-5.2 base model, [added by the neutral restatement:] and Z.ai attributes its benchmark gains [added by the neutral restatement:] to expanded post-training. Z.ai [added by the neutral restatement:] reports four specific score pairs: Z.ai Code Bench 20.9 to 31.4, Terminal-Bench 3.0 4.6 to 28.3, DeepSWE 46.2 to 66.9, AutomationBench 26.2 to 48.2.
Red-tinted words in the claim drifted from the evidence. Green-tinted words are what a neutral restatement needs.
The trace / claim to source
- Z.ai does state that GLM-5.3 uses the same base model as GLM-5.2 and that the gains come from scaled post-training. Bloomberg independently reports the same-base-model characterization, sourced to the company.
- Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9 appear verbatim in Z.ai's own developer documentation.
- AutomationBench 26.2 to 48.2 matches Z.ai's published table as reported by VentureBeat and other outlets.
- Z.ai Code Bench 20.9 to 31.4 matches Z.ai's High-effort figures and is consistent with the vendor's stated "50% improvement" headline.
- The claim's wording "almost entirely from post-training" is, if anything, weaker than what Z.ai asserts. The vendor says the gain is entirely from post-training. The claim does not overstate the vendor here.
- GLM-5.3 exists and was released on 2026-08-14 through the Z.ai API and GLM Coding Plan. This is not a fabricated announcement.
- Marketing as evidence: the claim presents four benchmark deltas as results, without stating that Z.ai produced all four itself inside its own harness. Only two rows in the roughly nineteen-row launch table were scored by named third parties, and none of the four cited here is among them. A reasonable reader takes these as measured facts about the model rather than as an interested party's self-report.
- Omitted qualifier: the Z.ai Code Bench pair 20.9 to 31.4 is the High-effort pair specifically. The same chart also shows a Max-effort pair of 23.4 to 34.5 and a Low-effort pair. The claim reports one point on a curve as if it were the benchmark result. It also omits that Z.ai Code Bench is private and cannot be rerun by anyone outside Z.ai.
- Benchmark cherry picking: the four rows selected are among the largest jumps in the table. The Terminal-Bench 3.0 figure rises from a near-floor 4.6, which mechanically produces the most dramatic multiple in the set. The same official table shows GLM-5.3 trailing GPT-5.6 Sol and Claude Fable 5 on Terminal-Bench 3.0 (28.3 versus 34.6 and 33.7), on DeepSWE v1.1 (66.9 versus 72.7 and 69.7), and badly on ExploitBench and ExploitGym. Nothing in the claim is false because of this, but the selection makes the release look like a clean sweep when the vendor's own chart does not.
- Harness mismatch: the DeepSWE baseline of 46.2 is Z.ai's own run under its own settings. The independent DeepSWE leaderboard listed GLM-5.2 at 44 plus or minus 2 using mini-swe-agent with standard settings. Same benchmark name, different evaluation record. The delta is therefore internally consistent within Z.ai's harness and is not directly comparable to the public board.
- Whether the base model is genuinely unchanged cannot be independently verified. The GLM-5.3 weights have not been released, so the "same base model" statement rests entirely on Z.ai's word. This is the load-bearing premise of the whole post and it is currently unauditable.
- No independent reproduction of any GLM-5.3 score exists as of 2026-08-15. Third parties have API access only, and the weights are held back for roughly two weeks pending safety hardening.
- I could not load the Z.ai launch blog page itself. The vendor figures here were read from Z.ai's developer documentation and from a mirror of the announcement footnotes, plus consistent reporting across named outlets. The AutomationBench and Z.ai Code Bench pairs specifically were confirmed through reporting rather than through a vendor page I retrieved directly.
- The claim writes all four figures as percentages. GDPval-AA v2 is an Elo score and ExploitGym is a task count, so the post's wider table mixes units, but the four cited figures do appear to be percentage metrics.
Z.ai released GLM-5.3 on 2026-08-14. Z.ai's own developer documentation states, in the vendor's words, that Terminal-Bench 3.0 increased from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5, and that GLM-5.3 improves by 50% over GLM-5.2 on Z.ai Code Bench. Multiple named outlets report the same table. VentureBeat records the AutomationBench pair as 26.2 to 48.2 and the same Terminal-Bench and DeepSWE pairs. Trade coverage records the Z.ai Code Bench pair as 20.9 to 31.4 at High effort, with a separate Max-effort pair of 23.4 to 34.5, and a token-efficiency claim of roughly 50,000 output tokens per task at High effort versus roughly 120,000 for Claude Opus 4.8 at 29.5%. On the architectural claim, Bloomberg reports that GLM-5.3 is built on the same roughly 700-billion-parameter base as its predecessor, and Z.ai's launch text is quoted by The Stack and Decrypt as saying that scaling post-training is all the company did for this release. Reported base specification is a 743B-parameter mixture-of-experts model with roughly 40B active parameters, unchanged from GLM-5.2. Every figure in the claim is Z.ai-run. The announcement footnotes, mirrored on OpenLM.ai, show that CyberGym, ExploitGym and the coding rows were evaluated by Z.ai inside Claude Code 2.1.207 at max reasoning effort, temperature 1.0, top_p 1.0, 128K max new tokens, single-run pass@1, unlimited timeout, with a domain whitelist. AutomationBench was run on v1.0.6 with a specific PR fix applied. Only GDPval-AA v2 (evaluated by Artificial Analysis) and Toolathlon Verified (official evaluation service) were scored outside Z.ai, which is two rows out of roughly nineteen. One provenance discrepancy is documented: the independent DeepSWE leaderboard, which runs all models on mini-swe-agent, listed GLM-5.2 at 44% plus or minus 2 at the time of the launch, not the 46.2 that appears in Z.ai's table. Z.ai's footnote also uses mini-swe-agent but specifies temperature 0.95, six-hour timeouts, 400K context and its own run. Weights are not public. Axios reports Z.ai is delaying the public release of the model weights for two weeks while it tests and strengthens safety controls, tied to the model's vulnerability-finding capability. As of the claim date no third party has run GLM-5.3 independently, because only API and Coding Plan access exist.
Complete reasoning
The reply is formatted for pasting into the thread where the claim is circulating.
Compact share page: ai.trueseeker.com/s/3013b5743a14/xDu4RbnHdNtyXNV4afqOdYHdH1S
Ask this case
Answers come only from the case file above; nothing is added.
Are the four benchmark numbers in the claim accurate?
Yes. Terminal-Bench 3.0 (4.6 to 28.3), DeepSWE v1.1 (46.2 to 66.9), AutomationBench (26.2 to 48.2), and Z.ai Code Bench (20.9 to 31.4) all match Z.ai's own developer documentation and launch materials, as confirmed by multiple outlets.
Who actually ran these benchmark tests?
Z.ai ran all four of these evaluations itself, inside its own harness. Only two rows out of roughly nineteen in the full launch table were scored by named third parties, and none of the four cited in this claim were among them.
Can anyone independently check these results by running GLM-5.3 themselves?
Not yet. As of the claim date, no third party has independently reproduced any GLM-5.3 result. The model weights are being withheld for about two weeks over cybersecurity concerns, so only API and Coding Plan access exist.
Is it true that GLM-5.3 uses the exact same base model as GLM-5.2?
Z.ai states this and Bloomberg reported the same characterization, but it cannot be independently verified since the weights have not been released. This claim rests entirely on the vendor's word.
Does the claim leave out anything important about how GLM-5.3 compares to other models?
Yes. Z.ai's own table shows GLM-5.3 trailing GPT-5.6 Sol and Claude Fable 5 on some of these same benchmarks, including Terminal-Bench 3.0 and DeepSWE, and performing badly on ExploitBench and ExploitGym. The claim does not mention this.