Quick answer: Claude Opus 4.5 is a significant checkpoint in Anthropic's Opus 4 update series, offering improved agentic capabilities and reasoning depth over the 4.1 base. It continues the $15/$75 pricing and 200K context window, targeting enterprise teams running complex, long-horizon workflows.
Claude Opus 4.5 represents a more substantial update in the Opus 4 numbered series than the 4.1 increment. Augment Code noted that Haiku 4.5 achieves 90% of Sonnet 4.5's performance — placing Sonnet 4.5 as the frontier model benchmark at the time of Haiku 4.5's release. This positions Opus 4.5 as the reasoning ceiling of its generation, handling the most demanding agentic tasks that smaller models in the series cannot reliably complete.
For Benchgen users, Opus 4.5 sits in the progression between Opus 4 and the Opus 4.6 release, which introduced a 1M-token context window in beta and raised SWE-bench Verified to 81.42%.
| Field | Value |
|---|---|
| Organization | Anthropic |
| Context window | 200,000 tokens |
| License | Proprietary |
| Modality | Multimodal (text and vision) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Anthropic | $15.00 | $75.00 |
Claude Opus 4.5 is a large language model developed by Anthropic.
| Model | ARC-AGI-v2 | GPQA-Diamond | CyberGym | License |
|---|---|---|---|---|
| Claude Opus 4.5 | 37.6% | 87.0% | 50.6% | Proprietary |
| Claude Opus 4.7 | — | — | — | Proprietary |
| Claude Opus 4 | — | — | — | Proprietary |
Claude Opus 4.5 GPQA-Diamond (87.0%) is strong. Opus 4.7 supersedes with 93.5% ARC-AGI.
Scores from Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.