Quick answer: GLM-5.2 is Zhipu AI's July 2026 frontier reasoning model, scoring 54.7% on Humanity's Last Exam, 99.2% on AIME 2026, 91.2% on GPQA Diamond, 77.0% on ARC-AGI, and 40.6% on Agents Last Exam. Its 54.7% HLE is exceptional — among the highest scores on this benchmark.
Where GLM-5.2 leads
Where it lags
Best for: Frontier reasoning and mathematics tasks; strong HLE-class academic research; teams in the Chinese AI ecosystem requiring maximum reasoning.
GLM-5.2 is Zhipu AI's July 2026 flagship model — part of the General Language Model (GLM) series that has been a cornerstone of Zhipu AI's model portfolio. The model shows an unusual characteristic: it achieves the highest HLE score (54.7%) of any model in this comparison, yet a lower ARC-AGI (77%) than frontier multimodal models.
This pattern suggests GLM-5.2 is specifically optimised for academic/research reasoning (HLE, GPQA, AIME) but less strong on general abstract reasoning patterns (ARC-AGI). The near-perfect AIME 2026 (99.2%) alongside 54.7% HLE makes it a prime choice for academic mathematics and science.
The 89.2% Global MMLU-Lite score reflects GLM's traditional strength in multilingual knowledge, particularly Chinese and other non-English languages.
| Field | Value |
|---|---|
| Organization | Zhipu AI |
| License | Proprietary (API only) |
| Release date | July 2026 |
| Knowledge cutoff | March 2026 |
| Modality | Text only |
Available via Zhipu AI open platform (bigmodel.cn). Refer to Zhipu pricing for current rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| Humanity's Last Exam | 54.7% | Benchgen evaluation | 2026-07 |
| AIME 2026 | 99.2% | Benchgen evaluation | 2026-07 |
| GPQA Diamond | 91.2% | Benchgen evaluation | 2026-07 |
| ARC-AGI | 77.0% | Benchgen evaluation | 2026-07 |
| Agents Last Exam | 40.6% | Benchgen evaluation | 2026-07 |
| Global MMLU-Lite | 89.2% | Benchgen evaluation | 2026-07 |
| Model | HLE | AIME 2026 | GPQA Diamond | Vision |
|---|---|---|---|---|
| GLM-5.2 | 54.7% | 99.2% | 91.2% | No |
| Gemini 3.1 Pro | 46.4% | 98.3% | 94.3% | Yes |
| GPT-5.5 | 41.4% | — | 93.6% | Yes |
GLM-5.2 leads on HLE (54.7% vs 46.4%) and matches Gemini 3.1 Pro on AIME (99.2% vs 98.3%) while lacking vision. For maximum academic reasoning without vision requirement: GLM-5.2. For multimodal + reasoning: Gemini 3.1 Pro.
Specs from Zhipu AI's official GLM-5.2 release (July 2026) and Benchgen evaluations. Last updated 2026-07-24.