Benchgen
Models/anthropic/

Claude Opus 4.5

DraftPublic

Model Details

Claude Opus 4.5

Organization Context Pricing License Modality

Quick answer: Claude Opus 4.5 is a significant checkpoint in Anthropic's Opus 4 update series, offering improved agentic capabilities and reasoning depth over the 4.1 base. It continues the $15/$75 pricing and 200K context window, targeting enterprise teams running complex, long-horizon workflows.

What Claude Opus 4.5 Is

Claude Opus 4.5 represents a more substantial update in the Opus 4 numbered series than the 4.1 increment. Augment Code noted that Haiku 4.5 achieves 90% of Sonnet 4.5's performance — placing Sonnet 4.5 as the frontier model benchmark at the time of Haiku 4.5's release. This positions Opus 4.5 as the reasoning ceiling of its generation, handling the most demanding agentic tasks that smaller models in the series cannot reliably complete.

For Benchgen users, Opus 4.5 sits in the progression between Opus 4 and the Opus 4.6 release, which introduced a 1M-token context window in beta and raised SWE-bench Verified to 81.42%.

Specifications

FieldValue
OrganizationAnthropic
Context window200,000 tokens
LicenseProprietary
ModalityMultimodal (text and vision)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Anthropic$15.00$75.00

Last updated 2026-06-19.

Claude Opus 4.5

Claude Opus 4.5 is a large language model developed by Anthropic.

Claude Opus 4.5 vs Alternatives

ModelARC-AGI-v2GPQA-DiamondCyberGymLicense
Claude Opus 4.537.6%87.0%50.6%Proprietary
Claude Opus 4.7Proprietary
Claude Opus 4Proprietary

Claude Opus 4.5 GPQA-Diamond (87.0%) is strong. Opus 4.7 supersedes with 93.5% ARC-AGI.

Frequently Asked Questions

What is Claude Opus 4.5? Anthropic's September 2025 model scoring 87.0% GPQA-Diamond, 50.6% CyberGym. Proprietary — strong scientific reasoning.

Scores from Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.