Benchgen
Models/anthropic/

Claude Opus 4

DraftPublic

Model Details

Claude Opus 4

Organization License Released Price

Quick answer: Claude Opus 4 is Anthropic's May 2025 flagship model, scoring 72.5% SWE-Bench Verified, 46.9% LiveCodeBench, 6.68% HLE, and 8.6% ARC-AGI v2. Priced at $15/$75 per MTok.

At a Glance

Where Claude Opus 4 leads

  • 72.5% SWE-Bench Verified — competitive real-world software engineering
  • Anthropic brand trust and safety research investment
  • Strong long-context reasoning
  • Extended thinking mode

Where it lags

  • 6.68% HLE — lower than newer frontier models
  • 8.6% ARC-AGI v2 — moderate on abstract reasoning
  • $15/$75 per MTok — high cost vs Claude Sonnet 4
  • Succeeded by Claude models in subsequent generations

Best for: Enterprise use cases requiring Anthropic's safety guarantees; complex agentic software engineering; long-context document analysis.

What Claude Opus 4 Is

Claude Opus 4 is Anthropic's May 2025 flagship — the most capable model in the Claude 4 generation at release. Claude models are known for their adherence to Constitutional AI principles, strong instruction following, and careful reasoning on complex tasks.

The 72.5% SWE-Bench Verified score reflects Opus 4's strength on real-world software engineering tasks — competitive at the time of release. The model supports extended thinking mode for longer chain-of-thought reasoning on complex tasks.

At $15/$75 per MTok, Opus 4 is positioned at the premium tier of Claude 4, with Claude Sonnet 4 and Haiku 4 offering lower-cost alternatives.

Specifications

FieldValue
OrganizationAnthropic
LicenseProprietary (API only)
Release dateMay 22, 2025
ModalityText only
Context window200K tokens
Output tokensUp to 32K

Pricing

TierPrice per MTok
Input$15.00
Output$75.00

Available via Anthropic API (api.anthropic.com) and Amazon Bedrock.

Public Benchmark Scores

BenchmarkScoreSourceDate
SWE-Bench Verified72.5%Benchgen evaluation2025-05
LiveCodeBench46.9%Benchgen evaluation2025-05
Humanity's Last Exam6.68%Benchgen evaluation2025-05
ARC-AGI v28.6%Benchgen evaluation2025-05
Shade Arena30.2%Benchgen evaluation2025-05

Claude Opus 4 vs Alternatives

ModelSWE-BenchLiveCodeBenchPrice (in/out)
Claude Opus 472.5%46.9%$15/$75
GPT-5 Codex74.5%Proprietary
HY378.0%Proprietary
DeepSeek-V3.2 Speciale73.1%MIT (open)

Claude Opus 4 at 72.5% SWE-Bench is competitive with GPT-5 Codex (74.5%) and HY3 (78%). For open-source SWE: DeepSeek-V3.2 Speciale (73.1%, MIT). For Anthropic ecosystem: Opus 4 with extended thinking is the primary choice.

Frequently Asked Questions

What is Claude Opus 4? Anthropic's May 2025 flagship model scoring 72.5% SWE-Bench Verified, 46.9% LiveCodeBench. $15/$75 per MTok. Top of the Claude 4 generation.
Is Claude Opus 4 worth $15/$75? For tasks requiring maximum Claude capability (complex reasoning, long-context, extended thinking), Opus 4 is justified. For most production coding/chat tasks, Claude Sonnet 4 offers better price/performance. Use Opus 4 when you need the flagship.

Specs from Anthropic's Claude Opus 4 release (May 2025) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.