Quick answer: Gemini 3.1 Pro is Google's March 2026 frontier multimodal model, scoring 98.0% on ARC-AGI, 98.3% on AIME 2026, 94.3% on GPQA Diamond, 46.4% on Humanity's Last Exam, and 85.9% on BrowseComp. It succeeds Gemini 2.5 Pro as Google's most capable model.
Where Gemini 3.1 Pro leads
Where it lags
Best for: Maximum quality multimodal AI; mathematics and science research; frontier agentic tasks requiring strongest Google capability.
Gemini 3.1 Pro (released March 2026) is Google's flagship frontier model, succeeding Gemini 2.5 Pro. The model demonstrates particularly exceptional performance on mathematical reasoning benchmarks — 98.3% AIME 2026 represents near-perfect performance on competition-level mathematics.
The 98.0% ARC-AGI score is also notable — this benchmark was designed to test abstract reasoning that AI models should find difficult. Gemini 3.1 Pro achieves near-ceiling performance on ARC-AGI, alongside 94.3% GPQA Diamond (graduate-level physics, biology, chemistry).
The 85.9% BrowseComp score (vs Gemini 2.5 Pro's strong showing on web research) suggests Gemini 3.1 Pro maintains Google's lead on web-grounded research tasks.
| Field | Value |
|---|---|
| Organization | |
| License | Proprietary (API only) |
| Release date | March 2026 |
| Knowledge cutoff | January 2026 |
| Modality | Multimodal |
Available via Google AI Studio and Vertex AI. Refer to Google's AI pricing page for current Gemini 3.1 Pro rates.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| ARC-AGI | 98.0% | Benchgen evaluation | 2026-03 |
| AIME 2026 | 98.3% | Benchgen evaluation | 2026-03 |
| GPQA Diamond | 94.3% | Benchgen evaluation | 2026-03 |
| Humanity's Last Exam | 46.4% | Benchgen evaluation | 2026-03 |
| BrowseComp | 85.9% | Benchgen evaluation | 2026-03 |
| CyberGym | 38.8% | Benchgen evaluation | 2026-03 |
| Model | ARC-AGI | GPQA Diamond | HLE | Vision |
|---|---|---|---|---|
| Gemini 3.1 Pro | 98.0% | 94.3% | 46.4% | Yes |
| Gemini 2.5 Pro | — | — | — | Yes |
| GPT-5.5 | 95.0% | 93.6% | 41.4% | Yes |
| GLM-5.2 | — | 91.2% | 54.7% | No |
Gemini 3.1 Pro leads on ARC-AGI (98% vs GPT-5.5's 95%) and AIME 2026 (98.3%). GLM-5.2 leads on HLE (54.7% vs 46.4%) but lacks vision.
Specs from Google's Gemini 3.1 Pro release (March 2026) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.