Benchgen
Models/openai/

GPT-5.4

DraftPublic

Model Details

GPT-5.4

Organization License Multimodal Released

Quick answer: GPT-5.4 is OpenAI's January 2026 frontier multimodal model, scoring 93.7% on ARC-AGI, 92.8% on GPQA Diamond, 82.7% on BrowseComp, 79.0% on CyberGym, and 39.8% on Humanity's Last Exam. It is the January 2026 point in GPT-5's incremental release cadence, superseded by GPT-5.5 in June 2026.

At a Glance

Where GPT-5.4 leads

  • 93.7% ARC-AGI — strong abstract reasoning
  • 92.8% GPQA Diamond — excellent graduate-level science
  • 82.7% BrowseComp — strong web research
  • 79.0% CyberGym — competitive cybersecurity
  • Full multimodal capability

Where it lags

  • Superseded by GPT-5.5 on most benchmarks
  • 39.8% HLE — below frontier models on Humanity's Last Exam
  • Proprietary

Best for: Production use cases validated on GPT-5.4; users migrating from GPT-5.x to this specific version.

What GPT-5.4 Is

GPT-5.4 (released January 2026) is one of OpenAI's incremental releases in the GPT-5 family. Following the GPT-4 family's cadence (gpt-4, gpt-4-turbo, gpt-4o), GPT-5 has been released as a versioned series with incremental improvements.

GPT-5.4's 93.7% ARC-AGI and 92.8% GPQA Diamond represent strong frontier performance. It was superseded by GPT-5.5 (June 2026) which improves on most benchmarks including ARC-AGI (95.0% vs 93.7%) and HLE (41.4% vs 39.8%).

Specifications

FieldValue
OrganizationOpenAI
LicenseProprietary (API only)
Release dateJanuary 2026
Knowledge cutoffSeptember 2025
ModalityMultimodal

Pricing

Available via OpenAI API. Refer to the OpenAI pricing page.

Public Benchmark Scores

BenchmarkScoreSourceDate
ARC-AGI93.7%Benchgen evaluation2026-01
GPQA Diamond92.8%Benchgen evaluation2026-01
BrowseComp82.7%Benchgen evaluation2026-01
CyberGym79.0%Benchgen evaluation2026-01
Humanity's Last Exam39.8%Benchgen evaluation2026-01
Agents Last Exam37.3%Benchgen evaluation2026-01

GPT-5.4 vs Alternatives

ModelARC-AGIGPQA DiamondHLEVersion
GPT-5.493.7%92.8%39.8%Jan 2026
GPT-5.595.0%93.6%41.4%Jun 2026
Gemini 3.1 Pro98.0%94.3%46.4%Mar 2026

GPT-5.5 strictly dominates GPT-5.4 on all available benchmarks. For new deployments, GPT-5.5 is the recommended GPT-5 series model.

Frequently Asked Questions

What is GPT-5.4? OpenAI's January 2026 frontier multimodal model scoring 93.7% ARC-AGI, 92.8% GPQA Diamond, and 39.8% HLE. Superseded by GPT-5.5 (June 2026).
Should I use GPT-5.4 or GPT-5.5? GPT-5.5 for new deployments — strictly better on all benchmarks (95.0% vs 93.7% ARC-AGI, 41.4% vs 39.8% HLE).

Specs from OpenAI's GPT-5.4 release (January 2026) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.