Quick answer: GLM-5.3 is Zhipu AI's August 2026 update to GLM-5.2 — same underlying base model, entirely scaled post-training. It scores 88.2% on Terminal-Bench 2.1, 28.3% on the new Terminal-Bench 3.0, 84.5% on CyberGym, 73.0% on Toolathlon Verified, and 1,769 Elo on GDPval-AA v2. Zhipu describes it as the most capable open-weights model for coding once weights ship (~2 weeks post-launch), and it shows a sharp, unexpected jump in cybersecurity/vulnerability-discovery capability versus GLM-5.2.
Where GLM-5.3 leads
Where it lags
thinking.type: "enabled" (with low/high/max effort), a breaking API change from GLM-5.2Best for: Coding agents and long-horizon agentic coding tasks (via Z.ai's Coding Plan/ZCode or the Claude Code/OpenCode integrations); security teams evaluating vulnerability-discovery capability; teams already on GLM-5.2 planning to upgrade once open weights ship.
GLM-5.3 is Zhipu AI's (Z.ai) August 2026 release, positioned explicitly as a post-training-only upgrade: "Scaling post-training is all we did for GLM-5.3." It reuses the same base model as GLM-5.2 and layers on more reinforcement-learning environments, more diverse long-horizon tasks, and more compute spent training on them — built on the same IndexShare (long-context), SAO (long-horizon RL), and slime (async large-scale training) stack introduced with GLM-5.2.
The headline result is a large jump in coding-agent capability: Zhipu reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, plus open-source state-of-the-art results on the public Terminal-Bench 3.0 and Agents' Last Exam benchmarks. Notably, GLM-5.3 achieves this while using fewer output tokens than GLM-5.2 at every effort level (e.g., 34.5% at ~75K tokens/task at Max effort, versus GLM-5.2's 23.4% at 96K).
A second, unplanned result is a sharp rise in cybersecurity capability. Zhipu added vulnerability-discovery data/environments to post-training expecting incremental gains, but found the model developed the ability to reason across multiple stages of a full exploitation chain — not just spot isolated flaws. On CyberGym it reaches 84.5% (from 77.2% for GLM-5.2), the best published result on that benchmark in Zhipu's own comparison. The gains are largest on exploitation-chain benchmarks further from the frontier's current strength (ExploitBench more than doubled, from 24.4% to 54.4%), though GLM-5.3 remains behind the closed frontier there. Zhipu also ran the model against real-world codebases with several security teams in China, surfacing 2,436 vulnerabilities (1,097 medium-to-high severity) across 269 open-source projects — some undiscovered for decades — now tracked via the public Z.ai Security Disclosure Ledger.
Weights are not yet released. Zhipu states they'll ship "in two weeks after launch, once safety evaluation and hardening are complete" — consistent with this tracker's original watchlist note (~Aug 28, 2026 target). Until then, GLM-5.3 is available only via the Z.ai API and GLM Coding Plan/ZCode.
| Field | Value |
|---|---|
| Organization | Zhipu AI (Z.ai) |
| License | Proprietary (API only) — open weights expected ~Aug 28, 2026 |
| Release date | August 14, 2026 |
| Knowledge cutoff | March 2026 (same base model/pretrain as GLM-5.2) |
| Modality | Text only |
| Base model | Same as GLM-5.2 (~743–753B parameters); all gains from post-training |
| Thinking | Always enabled — low/high/max reasoning effort; disabling thinking is no longer supported |
Available now via the Z.ai API and the GLM Coding Plan (points-based quota, 50% off-peak discount outside 14:00–18:00 UTC+8 on weekdays). Refer to Z.ai's pricing page for current per-token rates. Self-hosting will become possible once open weights ship (~Aug 28, 2026).
| Benchmark | Score | Source | Date |
|---|---|---|---|
| TerminalBench 2.1 | 88.2% | Z.ai launch blog | 2026-08 |
| Terminal-Bench 3.0 | 28.3% | Z.ai launch blog | 2026-08 |
| DeepSWE (v1.1) | 66.9% | Z.ai launch blog | 2026-08 |
| CyberGym | 84.5% | Z.ai launch blog | 2026-08 |
| Toolathlon-Verified | 73.0% | Z.ai launch blog | 2026-08 |
| GDPVal-AA v2 | 1,769 Elo | Z.ai launch blog | 2026-08 |
| NL2Repo-Bench | 58.0% | Z.ai launch blog | 2026-08 |
| ExploitGym (6h budget) | 130 | Z.ai launch blog | 2026-08 |
| ExploitBench | 54.4% | Z.ai launch blog | 2026-08 |
Additional benchmarks reported by Zhipu, not yet on Benchgen's leaderboard (either no existing benchmark page/model to compare against, or the reported number uses a variant/version/harness that isn't directly comparable to an existing page — see notes):
| Model | Terminal-Bench 2.1 | DeepSWE | CyberGym | Toolathlon Verified |
|---|---|---|---|---|
| GLM-5.3 | 88.2% | 66.9% | 84.5% | 73.0% |
| GLM-5.2 | 81.0% | 46.2% | 77.2% | 59.9% |
| Kimi K3 | 88.3% | 67.5% | — | 76.5% |
| Qwen3.8 Max | 86.6% | 56.6% | — | 72.5% |
GLM-5.3 is a substantial upgrade over its own predecessor across every listed benchmark (e.g., CyberGym 77.2%→84.5%, DeepSWE 46.2%→66.9%), and is competitive with (though not always ahead of) Kimi K3 and Qwen3.8 Max on coding/agentic tasks. Zhipu's own broader comparison (not independently reproduced here) also places GLM-5.3 behind top closed-frontier models (Claude-class, GPT-5.6 Sol) on the hardest cybersecurity exploitation benchmarks, while leading on CyberGym specifically.
Specs from Zhipu AI's official GLM-5.3 launch blog (z.ai/blog/glm-5.3, Aug 14, 2026) and Benchgen research. Last updated 2026-08-24.