Quick answer: Claude Opus 4 with extended thinking enabled is Anthropic's May 2025 frontier model at maximum reasoning depth. It scores 56.6% on LiveCodeBench and 10.72% on Humanity's Last Exam with thinking on. At $15/$75 per 1M tokens with a 200K context window, it targets the hardest coding and multi-step reasoning tasks.
Where Claude Opus 4 (Thinking) leads
Where it lags
Best for: Hardest multi-step reasoning tasks, research-grade code generation, and production pipelines where auditability of reasoning steps is required.
Claude Opus 4 is Anthropic's May 2025 flagship model — the highest-capability tier in the Claude 4 family. The "thinking" configuration enables extended thinking mode, which gives the model a scratchpad for internal reasoning before producing its response. This improves accuracy on multi-step problems at the cost of higher latency and more output tokens.
The model's 56.6% LiveCodeBench score reflects strong coding capability when reasoning is applied. Combined with the 10.72% HLE score, it represents Anthropic's strongest general reasoning capability at the Opus 4 generation.
For cost-sensitive workloads, Claude Sonnet 4 (Thinking) provides similar LiveCodeBench performance (55.9%) at significantly lower cost ($3/$15 per 1M). Claude Opus 4 (Thinking) is the correct choice when maximum accuracy on complex tasks is the primary requirement regardless of cost.
| Field | Value |
|---|---|
| Organization | Anthropic |
| Parameters | Undisclosed |
| Context window | 200,000 tokens |
| Max output | 32,000 tokens (including thinking tokens) |
| License | Proprietary (API only) |
| Release date | May 22, 2025 |
| Knowledge cutoff | March 2025 |
| Thinking mode | Extended (visible reasoning) |
| Modality | Text + Vision (multimodal) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Anthropic API | $15.00 | $75.00 |
Thinking tokens billed as output tokens. Prompt caching available. Pricing per Anthropic pricing page.
Claude Opus 4 (Thinking) has a 200,000-token context window — roughly 150 pages of text in a single request.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| LiveCodeBench | 56.6% | Benchgen evaluation | 2025-07 |
| Humanity's Last Exam | 10.72% | Benchgen evaluation | 2025-07 |
| Model | LiveCodeBench | HLE | Thinking | Price (in/out per 1M) |
|---|---|---|---|---|
| Claude Opus 4 (Thinking) | 56.6% | 10.72% | Yes | $15 / $75 |
| Claude Sonnet 4 (Thinking) | 55.9% | 7.76% | Yes | $3 / $15 |
| o3 | 75.8% | 20.32% | Yes | $10 / $40 |
| o4-mini (high) | 80.2% | 18.08% | Yes | $1.10 / $4.40 |
| Gemini 2.5 Pro | 73.6% | 21.64% | Yes | $1.25 / $10 |
Claude Opus 4 Thinking's LiveCodeBench (56.6%) is significantly below o3 (75.8%) and o4-mini-high (80.2%). For coding-intensive workloads, OpenAI o-series or Gemini 2.5 Pro provide better results. Claude Opus 4 Thinking is differentiated by Anthropic's safety properties and visible chain-of-thought reasoning.
Specs and scores from Anthropic's official Claude 4 announcement (May 2025) and Benchgen evaluations. Pricing cited to the Anthropic pricing page. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.