Benchgen
Models/openai/

GPT-5.5

DraftPublic

Model Details

GPT-5.5

Organization License Multimodal Released

Quick answer: GPT-5.5 is OpenAI's June 2026 frontier multimodal model, scoring 95.0% on ARC-AGI, 93.6% on GPQA Diamond, 84.4% on BrowseComp, 81.8% on CyberGym, 47.9% on Agents Last Exam, and 41.4% on Humanity's Last Exam. The current top-tier GPT-5 series model.

At a Glance

Where GPT-5.5 leads

  • 95.0% ARC-AGI — near-ceiling abstract reasoning (+1.3pp vs GPT-5.4)
  • 93.6% GPQA Diamond — state-of-the-art graduate science
  • 81.8% CyberGym — leading cybersecurity benchmark
  • 84.4% BrowseComp — strong web research
  • 47.9% Agents Last Exam — strong agentic task performance
  • Full multimodal capability

Where it lags

  • 41.4% HLE — below Gemini 3.1 Pro (46.4%) and GLM-5.2 (54.7%)
  • ARC-AGI still below Gemini 3.1 Pro (98.0% vs 95.0%)
  • Proprietary

Best for: Maximum OpenAI capability; agentic tasks requiring the strongest GPT-5 model; multimodal frontier use cases.

What GPT-5.5 Is

GPT-5.5 (released June 2026) is OpenAI's most capable model at this release date — the June 2026 increment of the GPT-5 series. It improves on GPT-5.4 across all available benchmarks and adds stronger agentic performance (47.9% Agents Last Exam vs GPT-5.4's 37.3%).

The 81.8% CyberGym score is notable — significantly above GPT-5.4's 79.0% and placing it among the strongest models on security evaluation. The BrowseComp improvement (84.4% vs 82.7%) also indicates stronger web research capability.

Specifications

FieldValue
OrganizationOpenAI
LicenseProprietary (API only)
Release dateJune 2026
Knowledge cutoffSeptember 2025
ModalityMultimodal

Pricing

Available via OpenAI API. Refer to the OpenAI pricing page for current GPT-5.5 rates.

Public Benchmark Scores

BenchmarkScoreSourceDate
ARC-AGI95.0%Benchgen evaluation2026-06
GPQA Diamond93.6%Benchgen evaluation2026-06
BrowseComp84.4%Benchgen evaluation2026-06
CyberGym81.8%Benchgen evaluation2026-06
Agents Last Exam47.9%Benchgen evaluation2026-06
Humanity's Last Exam41.4%Benchgen evaluation2026-06

GPT-5.5 vs Alternatives

ModelARC-AGIGPQA DiamondHLECyberGym
GPT-5.595.0%93.6%41.4%81.8%
GPT-5.493.7%92.8%39.8%79.0%
Gemini 3.1 Pro98.0%94.3%46.4%38.8%
GLM-5.291.2%54.7%

GPT-5.5 vs Gemini 3.1 Pro: GPT-5.5 leads on CyberGym (81.8% vs 38.8%); Gemini leads on ARC-AGI (98% vs 95%) and HLE (46.4% vs 41.4%). Choice depends on task: GPT-5.5 for security/coding, Gemini 3.1 for math/science.

Frequently Asked Questions

What is GPT-5.5? OpenAI's June 2026 frontier multimodal model scoring 95.0% ARC-AGI, 93.6% GPQA Diamond, 81.8% CyberGym, and 41.4% HLE. The current top-tier GPT-5 series model.
How does GPT-5.5 compare to Gemini 3.1 Pro? Both are frontier multimodal models. Gemini 3.1 Pro leads on ARC-AGI (98% vs 95%), AIME 2026 (98.3%), and HLE (46.4% vs 41.4%). GPT-5.5 leads on CyberGym (81.8% vs 38.8%). Choice depends on task requirements.

Specs from OpenAI's GPT-5.5 release (June 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.