Benchgen
Models/anthropic/

Claude Sonnet 4 (Thinking)

DraftPublic

Model Details

Claude Sonnet 4 (Thinking)

Organization Context Pricing License Modality Released

Quick answer: Claude Sonnet 4 with extended thinking enabled scores 55.9% on LiveCodeBench and 7.76% on Humanity's Last Exam. It is the mid-tier option in the Claude 4 thinking family at $3/$15 per 1M tokens — nearly matching Claude Opus 4 Thinking on coding (55.9% vs 56.6%) at 80% lower cost, making it the highest-value extended-thinking model in the Claude 4 line.

At a Glance

Where Claude Sonnet 4 (Thinking) leads

  • 55.9% LiveCodeBench — nearly matches Opus 4 Thinking (56.6%) at 80% lower cost
  • $3/$15 per 1M — best price/performance for extended-thinking in Claude 4
  • Extended thinking: visible chain-of-thought reasoning
  • 200K context window
  • Computer use capability

Where it lags

  • 7.76% HLE — below Opus 4 Thinking (10.72%) on expert-level tasks
  • Trailing o4-mini-high (80.2% LiveCodeBench) and o3 (75.8%) on coding
  • $3/$15 higher than o4-mini ($1.10/$4.40) for lower coding performance

Best for: Extended-thinking coding and reasoning tasks at Sonnet-tier pricing; the recommended Claude 4 thinking model for the majority of use cases.

What Claude Sonnet 4 (Thinking) Is

Claude Sonnet 4 (Thinking) is the mid-tier extended-thinking model in Anthropic's May 2025 Claude 4 family. With thinking enabled, it achieves 55.9% LiveCodeBench — within 0.7 percentage points of Claude Opus 4 Thinking (56.6%), while costing 5× less per token. This makes it the highest-value extended-thinking model in the Claude 4 lineup.

The extended thinking mode provides a visible reasoning scratchpad, allowing users to inspect the model's reasoning process. This is valuable for debugging, verification, and tasks requiring transparent multi-step problem solving.

Claude Sonnet 4 Thinking replaced Claude 3.7 Sonnet as the recommended mid-tier extended-thinking model with the May 2025 Claude 4 launch. For most coding and reasoning tasks that require thinking, Claude Sonnet 4 Thinking is the recommended choice over Opus 4 Thinking.

Specifications

FieldValue
OrganizationAnthropic
ParametersUndisclosed
Context window200,000 tokens
Max output32,000 tokens (including thinking tokens)
LicenseProprietary (API only)
Release dateMay 22, 2025
Knowledge cutoffMarch 2025
Thinking modeExtended (visible reasoning)
ModalityText + Vision (multimodal)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Anthropic API$3.00$15.00

Thinking tokens billed as output tokens. Prompt caching available. Pricing per Anthropic pricing page.

Public Benchmark Scores

BenchmarkScoreSourceDate
LiveCodeBench55.9%Benchgen evaluation2025-07
Humanity's Last Exam7.76%Benchgen evaluation2025-07

Claude Sonnet 4 (Thinking) vs Alternatives

ModelLiveCodeBenchHLEThinkingPrice (in/out per 1M)
Claude Sonnet 4 (Thinking)55.9%7.76%Yes$3 / $15
Claude Opus 4 (Thinking)56.6%10.72%Yes$15 / $75
Claude 3.7 Sonnet8.04%Yes$3 / $15
o4-mini (high)80.2%18.08%Yes$1.10 / $4.40

Claude Sonnet 4 Thinking nearly matches Opus 4 Thinking on LiveCodeBench at $3/$15 — making it the better value for most coding tasks. o4-mini-high dominates on coding benchmarks at lower cost, making it the preferred choice for purely coding-intensive workloads. Claude Sonnet 4 Thinking is differentiated by Anthropic's safety properties and visible thinking.

Frequently Asked Questions

What is Claude Sonnet 4 Thinking? Claude Sonnet 4 (Thinking) is Anthropic's May 2025 mid-tier model with extended thinking enabled, scoring 55.9% LiveCodeBench and 7.76% HLE. Priced at $3/$15 per 1M tokens.
Should I use Claude Sonnet 4 Thinking or Opus 4 Thinking? For coding tasks, Claude Sonnet 4 Thinking is the better value: 55.9% vs 56.6% LiveCodeBench at 80% lower cost. Use Claude Opus 4 Thinking only when the HLE gap (10.72% vs 7.76%) matters for your specific expert-knowledge task.
How much does Claude Sonnet 4 (Thinking) cost? $3.00 per 1M input tokens and $15.00 per 1M output tokens. Thinking tokens count as output tokens.

Specs and scores from Anthropic's official Claude 4 announcement (May 2025) and Benchgen evaluations. Pricing cited to the Anthropic pricing page. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.