Benchgen
Models/google/

Gemini 3.1 Pro

DraftPublic

Model Details

Gemini 3.1 Pro

Organization License Multimodal Released

Quick answer: Gemini 3.1 Pro is Google's March 2026 frontier multimodal model, scoring 98.0% on ARC-AGI, 98.3% on AIME 2026, 94.3% on GPQA Diamond, 46.4% on Humanity's Last Exam, and 85.9% on BrowseComp. It succeeds Gemini 2.5 Pro as Google's most capable model.

At a Glance

Where Gemini 3.1 Pro leads

  • 98.0% ARC-AGI — near-ceiling abstract reasoning
  • 98.3% AIME 2026 — near-perfect competition mathematics
  • 94.3% GPQA Diamond — exceptional graduate-level science
  • 85.9% BrowseComp — strong web research capability
  • 46.4% Humanity's Last Exam — top-tier frontier performance
  • Full multimodal capability
  • January 2026 knowledge cutoff

Where it lags

  • 38.8% CyberGym — below some security-focused models
  • 33% AA Omniscience — limited niche performance
  • Proprietary: no open weights

Best for: Maximum quality multimodal AI; mathematics and science research; frontier agentic tasks requiring strongest Google capability.

What Gemini 3.1 Pro Is

Gemini 3.1 Pro (released March 2026) is Google's flagship frontier model, succeeding Gemini 2.5 Pro. The model demonstrates particularly exceptional performance on mathematical reasoning benchmarks — 98.3% AIME 2026 represents near-perfect performance on competition-level mathematics.

The 98.0% ARC-AGI score is also notable — this benchmark was designed to test abstract reasoning that AI models should find difficult. Gemini 3.1 Pro achieves near-ceiling performance on ARC-AGI, alongside 94.3% GPQA Diamond (graduate-level physics, biology, chemistry).

The 85.9% BrowseComp score (vs Gemini 2.5 Pro's strong showing on web research) suggests Gemini 3.1 Pro maintains Google's lead on web-grounded research tasks.

Specifications

FieldValue
OrganizationGoogle
LicenseProprietary (API only)
Release dateMarch 2026
Knowledge cutoffJanuary 2026
ModalityMultimodal

Pricing

Available via Google AI Studio and Vertex AI. Refer to Google's AI pricing page for current Gemini 3.1 Pro rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
ARC-AGI98.0%Benchgen evaluation2026-03
AIME 202698.3%Benchgen evaluation2026-03
GPQA Diamond94.3%Benchgen evaluation2026-03
Humanity's Last Exam46.4%Benchgen evaluation2026-03
BrowseComp85.9%Benchgen evaluation2026-03
CyberGym38.8%Benchgen evaluation2026-03

Gemini 3.1 Pro vs Alternatives

ModelARC-AGIGPQA DiamondHLEVision
Gemini 3.1 Pro98.0%94.3%46.4%Yes
Gemini 2.5 ProYes
GPT-5.595.0%93.6%41.4%Yes
GLM-5.291.2%54.7%No

Gemini 3.1 Pro leads on ARC-AGI (98% vs GPT-5.5's 95%) and AIME 2026 (98.3%). GLM-5.2 leads on HLE (54.7% vs 46.4%) but lacks vision.

Frequently Asked Questions

What is Gemini 3.1 Pro? Google's March 2026 frontier multimodal model scoring 98.0% ARC-AGI, 98.3% AIME 2026, 94.3% GPQA Diamond, and 46.4% HLE. Google's strongest model.
How does Gemini 3.1 Pro compare to Gemini 2.5 Pro? Gemini 3.1 Pro is the successor — significantly improved on frontier benchmarks (ARC-AGI, AIME 2026, GPQA Diamond) with a more recent January 2026 knowledge cutoff.

Specs from Google's Gemini 3.1 Pro release (March 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.