Benchgen
Models/zhipu-ai/

GLM 5.2

DraftPublic

Model Details

GLM-5.2

Organization License Released

Quick answer: GLM-5.2 is Zhipu AI's July 2026 frontier reasoning model, scoring 54.7% on Humanity's Last Exam, 99.2% on AIME 2026, 91.2% on GPQA Diamond, 77.0% on ARC-AGI, and 40.6% on Agents Last Exam. Its 54.7% HLE is exceptional — among the highest scores on this benchmark.

At a Glance

Where GLM-5.2 leads

  • 54.7% HLE — among the highest Humanity's Last Exam scores publicly reported
  • 99.2% AIME 2026 — near-perfect competition mathematics
  • 91.2% GPQA Diamond — excellent graduate-level science
  • 89.2% Global MMLU-Lite — strong multilingual academic knowledge
  • Strong Chinese-language and STEM performance

Where it lags

  • 77.0% ARC-AGI — below Gemini 3.1 Pro (98%) and GPT-5.5 (95%)
  • Text-only; no vision (unlike Gemini/GPT-5.5)
  • Proprietary, API-only
  • Primarily Chinese ecosystem model

Best for: Frontier reasoning and mathematics tasks; strong HLE-class academic research; teams in the Chinese AI ecosystem requiring maximum reasoning.

What GLM-5.2 Is

GLM-5.2 is Zhipu AI's July 2026 flagship model — part of the General Language Model (GLM) series that has been a cornerstone of Zhipu AI's model portfolio. The model shows an unusual characteristic: it achieves the highest HLE score (54.7%) of any model in this comparison, yet a lower ARC-AGI (77%) than frontier multimodal models.

This pattern suggests GLM-5.2 is specifically optimised for academic/research reasoning (HLE, GPQA, AIME) but less strong on general abstract reasoning patterns (ARC-AGI). The near-perfect AIME 2026 (99.2%) alongside 54.7% HLE makes it a prime choice for academic mathematics and science.

The 89.2% Global MMLU-Lite score reflects GLM's traditional strength in multilingual knowledge, particularly Chinese and other non-English languages.

Specifications

FieldValue
OrganizationZhipu AI
LicenseProprietary (API only)
Release dateJuly 2026
Knowledge cutoffMarch 2026
ModalityText only

Pricing

Available via Zhipu AI open platform (bigmodel.cn). Refer to Zhipu pricing for current rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
Humanity's Last Exam54.7%Benchgen evaluation2026-07
AIME 202699.2%Benchgen evaluation2026-07
GPQA Diamond91.2%Benchgen evaluation2026-07
ARC-AGI77.0%Benchgen evaluation2026-07
Agents Last Exam40.6%Benchgen evaluation2026-07
Global MMLU-Lite89.2%Benchgen evaluation2026-07

GLM-5.2 vs Alternatives

ModelHLEAIME 2026GPQA DiamondVision
GLM-5.254.7%99.2%91.2%No
Gemini 3.1 Pro46.4%98.3%94.3%Yes
GPT-5.541.4%93.6%Yes

GLM-5.2 leads on HLE (54.7% vs 46.4%) and matches Gemini 3.1 Pro on AIME (99.2% vs 98.3%) while lacking vision. For maximum academic reasoning without vision requirement: GLM-5.2. For multimodal + reasoning: Gemini 3.1 Pro.

Frequently Asked Questions

What is GLM-5.2? Zhipu AI's July 2026 frontier reasoning model scoring 54.7% HLE, 99.2% AIME 2026, 91.2% GPQA Diamond, and 77.0% ARC-AGI. Highest publicly reported HLE score.
How does GLM-5.2 compare to GPT-5.5 and Gemini 3.1 Pro on benchmarks? GLM-5.2 leads on HLE (54.7% vs 46.4% Gemini, 41.4% GPT-5.5) and AIME 2026. Gemini 3.1 Pro leads on ARC-AGI (98% vs 77%) and adds vision. Choose based on task: academic reasoning → GLM-5.2; multimodal + abstract reasoning → Gemini 3.1 Pro.

Specs from Zhipu AI's official GLM-5.2 release (July 2026) and Benchgen evaluations. Last updated 2026-07-24.