Quick answer: Claude Opus 4 is Anthropic's May 2025 flagship model, scoring 72.5% SWE-Bench Verified, 46.9% LiveCodeBench, 6.68% HLE, and 8.6% ARC-AGI v2. Priced at $15/$75 per MTok.
Where Claude Opus 4 leads
Where it lags
Best for: Enterprise use cases requiring Anthropic's safety guarantees; complex agentic software engineering; long-context document analysis.
Claude Opus 4 is Anthropic's May 2025 flagship — the most capable model in the Claude 4 generation at release. Claude models are known for their adherence to Constitutional AI principles, strong instruction following, and careful reasoning on complex tasks.
The 72.5% SWE-Bench Verified score reflects Opus 4's strength on real-world software engineering tasks — competitive at the time of release. The model supports extended thinking mode for longer chain-of-thought reasoning on complex tasks.
At $15/$75 per MTok, Opus 4 is positioned at the premium tier of Claude 4, with Claude Sonnet 4 and Haiku 4 offering lower-cost alternatives.
| Field | Value |
|---|---|
| Organization | Anthropic |
| License | Proprietary (API only) |
| Release date | May 22, 2025 |
| Modality | Text only |
| Context window | 200K tokens |
| Output tokens | Up to 32K |
| Tier | Price per MTok |
|---|---|
| Input | $15.00 |
| Output | $75.00 |
Available via Anthropic API (api.anthropic.com) and Amazon Bedrock.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| SWE-Bench Verified | 72.5% | Benchgen evaluation | 2025-05 |
| LiveCodeBench | 46.9% | Benchgen evaluation | 2025-05 |
| Humanity's Last Exam | 6.68% | Benchgen evaluation | 2025-05 |
| ARC-AGI v2 | 8.6% | Benchgen evaluation | 2025-05 |
| Shade Arena | 30.2% | Benchgen evaluation | 2025-05 |
| Model | SWE-Bench | LiveCodeBench | Price (in/out) |
|---|---|---|---|
| Claude Opus 4 | 72.5% | 46.9% | $15/$75 |
| GPT-5 Codex | 74.5% | — | Proprietary |
| HY3 | 78.0% | — | Proprietary |
| DeepSeek-V3.2 Speciale | 73.1% | — | MIT (open) |
Claude Opus 4 at 72.5% SWE-Bench is competitive with GPT-5 Codex (74.5%) and HY3 (78%). For open-source SWE: DeepSeek-V3.2 Speciale (73.1%, MIT). For Anthropic ecosystem: Opus 4 with extended thinking is the primary choice.
Specs from Anthropic's Claude Opus 4 release (May 2025) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.