Quick answer: GPT-5.2 Pro 2025-12-11 is OpenAI's December 11, 2025 model scoring 93.2% GPQA Diamond, 90.5% ARC-AGI, 77.9% BrowseComp, and 36.6% HLE. Top GPQA Diamond score in its release window.
Where GPT-5.2 Pro leads
Where it lags
Best for: Applications requiring top GPQA Diamond science reasoning; ARC-AGI visual reasoning; BrowseComp web research; OpenAI ecosystem at December 2025.
GPT-5.2 Pro 2025-12-11 is the December 11, 2025 checkpoint of OpenAI's GPT-5.2 Pro model. The 2025-12-11 suffix identifies a specific deployment version. The 93.2% GPQA Diamond is the standout score — one of the highest publicly reported GPQA Diamond scores at the time of release.
The 90.5% ARC-AGI further shows strong abstract pattern recognition. Combined with 77.9% BrowseComp, this is a well-rounded model for science + research tasks.
| Field | Value |
|---|---|
| Organization | OpenAI |
| License | Proprietary (API only) |
| Release date | December 11, 2025 |
| Modality | Text only |
Available via OpenAI API. Check pricing at platform.openai.com.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| GPQA Diamond | 93.2% | Benchgen evaluation | 2025-12 |
| ARC-AGI | 90.5% | Benchgen evaluation | 2025-12 |
| BrowseComp | 77.9% | Benchgen evaluation | 2025-12 |
| Humanity's Last Exam | 36.6% | Benchgen evaluation | 2025-12 |
| Model | GPQA Diamond | ARC-AGI | BrowseComp | HLE |
|---|---|---|---|---|
| GPT-5.2 Pro 2025-12-11 | 93.2% | 90.5% | 77.9% | 36.6% |
| Meta Muse Spark | 89.5% | — | — | 40.6% |
| Claude Opus 4 | 90.1% | — | — | — |
| Kimi K2.6 | 90.5% | — | 83.2% | 36.4% |
GPT-5.2 Pro leads on GPQA Diamond (93.2%) over all listed models at December 2025. For BrowseComp: Kimi K2.6 (83.2% vs 77.9%). For higher HLE: Meta Muse Spark (40.6% vs 36.6%).
Specs from OpenAI's GPT-5.2 Pro 2025-12-11 release and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.