Benchgen
Models/anthropic/

Claude Opus 4 (Thinking)

DraftPublic

Model Details

Claude Opus 4 (Thinking)

Organization Context Pricing License Modality Released

Quick answer: Claude Opus 4 with extended thinking enabled is Anthropic's May 2025 frontier model at maximum reasoning depth. It scores 56.6% on LiveCodeBench and 10.72% on Humanity's Last Exam with thinking on. At $15/$75 per 1M tokens with a 200K context window, it targets the hardest coding and multi-step reasoning tasks.

At a Glance

Where Claude Opus 4 (Thinking) leads

  • 56.6% LiveCodeBench — strong competitive programming with extended reasoning
  • 10.72% Humanity's Last Exam — best in the Claude 4 family on expert tasks
  • Extended thinking: visible chain-of-thought reasoning for debugging and auditability
  • Computer use capability
  • 200K context window

Where it lags

  • $15/$75 per 1M tokens — highest cost in the Claude 4 family
  • 10.72% HLE trails o3 (20.32%), Gemini 2.5 Pro (21.64%), and GPT-5.6 Sol
  • No open weights

Best for: Hardest multi-step reasoning tasks, research-grade code generation, and production pipelines where auditability of reasoning steps is required.

What Claude Opus 4 (Thinking) Is

Claude Opus 4 is Anthropic's May 2025 flagship model — the highest-capability tier in the Claude 4 family. The "thinking" configuration enables extended thinking mode, which gives the model a scratchpad for internal reasoning before producing its response. This improves accuracy on multi-step problems at the cost of higher latency and more output tokens.

The model's 56.6% LiveCodeBench score reflects strong coding capability when reasoning is applied. Combined with the 10.72% HLE score, it represents Anthropic's strongest general reasoning capability at the Opus 4 generation.

For cost-sensitive workloads, Claude Sonnet 4 (Thinking) provides similar LiveCodeBench performance (55.9%) at significantly lower cost ($3/$15 per 1M). Claude Opus 4 (Thinking) is the correct choice when maximum accuracy on complex tasks is the primary requirement regardless of cost.

Specifications

FieldValue
OrganizationAnthropic
ParametersUndisclosed
Context window200,000 tokens
Max output32,000 tokens (including thinking tokens)
LicenseProprietary (API only)
Release dateMay 22, 2025
Knowledge cutoffMarch 2025
Thinking modeExtended (visible reasoning)
ModalityText + Vision (multimodal)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Anthropic API$15.00$75.00

Thinking tokens billed as output tokens. Prompt caching available. Pricing per Anthropic pricing page.

Context Window

Claude Opus 4 (Thinking) has a 200,000-token context window — roughly 150 pages of text in a single request.

Public Benchmark Scores

BenchmarkScoreSourceDate
LiveCodeBench56.6%Benchgen evaluation2025-07
Humanity's Last Exam10.72%Benchgen evaluation2025-07

Claude Opus 4 (Thinking) vs Alternatives

ModelLiveCodeBenchHLEThinkingPrice (in/out per 1M)
Claude Opus 4 (Thinking)56.6%10.72%Yes$15 / $75
Claude Sonnet 4 (Thinking)55.9%7.76%Yes$3 / $15
o375.8%20.32%Yes$10 / $40
o4-mini (high)80.2%18.08%Yes$1.10 / $4.40
Gemini 2.5 Pro73.6%21.64%Yes$1.25 / $10

Claude Opus 4 Thinking's LiveCodeBench (56.6%) is significantly below o3 (75.8%) and o4-mini-high (80.2%). For coding-intensive workloads, OpenAI o-series or Gemini 2.5 Pro provide better results. Claude Opus 4 Thinking is differentiated by Anthropic's safety properties and visible chain-of-thought reasoning.

Frequently Asked Questions

What is Claude Opus 4 Thinking? Claude Opus 4 (Thinking) is Anthropic's May 2025 flagship model with extended thinking mode enabled, scoring 56.6% LiveCodeBench and 10.72% HLE. Priced at $15/$75 per 1M tokens.
What is extended thinking in Claude Opus 4? Extended thinking enables a visible reasoning scratchpad before the final response. The model uses this to break down complex problems step-by-step. Thinking tokens are billed as output tokens.
Should I use Claude Opus 4 or Sonnet 4 Thinking? Claude Sonnet 4 (Thinking) scores 55.9% LiveCodeBench (vs 56.6%) and 7.76% HLE (vs 10.72%) at $3/$15 per 1M instead of $15/$75. For most workflows, Claude Sonnet 4 Thinking provides nearly identical coding performance at 80% lower cost. Use Opus 4 Thinking for the most demanding accuracy-critical tasks.

Specs and scores from Anthropic's official Claude 4 announcement (May 2025) and Benchgen evaluations. Pricing cited to the Anthropic pricing page. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.