Benchgen
Models/openai/

GPT-5.1

DraftPublic

Model Details

GPT-5.1

Organization Context Pricing License Modality

Quick answer: GPT-5.1 is OpenAI's iterative update to GPT-5, continuing the August 2025 flagship with improvements to instruction following, agentic coding reliability, and long-context reasoning. It maintains the 400K-token context window and $1.25/$10 per million token pricing of GPT-5.

What GPT-5.1 Is

GPT-5.1 follows OpenAI's versioned checkpoint pattern — periodic updates to the GPT-5 family that improve specific capability areas without a full architectural change. The .1 update focuses on the agentic and tool-use capabilities that GPT-5 introduced, with refinements that reduce error rates in multi-step workflows and improve adherence to complex system prompts.

For Benchgen users, GPT-5.1 represents the same fundamental reasoning architecture as GPT-5, with the same router system (fast model + thinking model) and pricing, but with measurable improvements in the tasks most relevant to agent evaluation: coding precision, tool-call reliability, and multi-turn context retention.

Specifications

FieldValue
OrganizationOpenAI
Context window400,000 tokens
Max output128,000 tokens
LicenseProprietary
Base seriesGPT-5 (August 2025)
ModalityMultimodal (text and vision)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
OpenAI$1.25$10.00

GPT-5.1 vs Alternatives

ModelBrowseCompGPQA-DiamondSimpleQALicense
GPT-5.190.0%88.1%45.6%Proprietary
GPT-5.1 Thinking90.0%88.1%Proprietary
GPT-5.1 Instant90.0%88.1%Proprietary

GPT-5.1 base model vs Thinking vs Instant: same BrowseComp + GPQA-Diamond across all three. Differentiated by HLE and speed.

Frequently Asked Questions

What is GPT-5.1? OpenAI's January 2026 model scoring 90.0% BrowseComp, 88.1% GPQA-Diamond, 45.6% SimpleQA. Proprietary.

Scores from Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.