Quick answer: Claude Sonnet 4 with extended thinking enabled scores 55.9% on LiveCodeBench and 7.76% on Humanity's Last Exam. It is the mid-tier option in the Claude 4 thinking family at $3/$15 per 1M tokens — nearly matching Claude Opus 4 Thinking on coding (55.9% vs 56.6%) at 80% lower cost, making it the highest-value extended-thinking model in the Claude 4 line.
Where Claude Sonnet 4 (Thinking) leads
Where it lags
Best for: Extended-thinking coding and reasoning tasks at Sonnet-tier pricing; the recommended Claude 4 thinking model for the majority of use cases.
Claude Sonnet 4 (Thinking) is the mid-tier extended-thinking model in Anthropic's May 2025 Claude 4 family. With thinking enabled, it achieves 55.9% LiveCodeBench — within 0.7 percentage points of Claude Opus 4 Thinking (56.6%), while costing 5× less per token. This makes it the highest-value extended-thinking model in the Claude 4 lineup.
The extended thinking mode provides a visible reasoning scratchpad, allowing users to inspect the model's reasoning process. This is valuable for debugging, verification, and tasks requiring transparent multi-step problem solving.
Claude Sonnet 4 Thinking replaced Claude 3.7 Sonnet as the recommended mid-tier extended-thinking model with the May 2025 Claude 4 launch. For most coding and reasoning tasks that require thinking, Claude Sonnet 4 Thinking is the recommended choice over Opus 4 Thinking.
| Field | Value |
|---|---|
| Organization | Anthropic |
| Parameters | Undisclosed |
| Context window | 200,000 tokens |
| Max output | 32,000 tokens (including thinking tokens) |
| License | Proprietary (API only) |
| Release date | May 22, 2025 |
| Knowledge cutoff | March 2025 |
| Thinking mode | Extended (visible reasoning) |
| Modality | Text + Vision (multimodal) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Anthropic API | $3.00 | $15.00 |
Thinking tokens billed as output tokens. Prompt caching available. Pricing per Anthropic pricing page.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| LiveCodeBench | 55.9% | Benchgen evaluation | 2025-07 |
| Humanity's Last Exam | 7.76% | Benchgen evaluation | 2025-07 |
| Model | LiveCodeBench | HLE | Thinking | Price (in/out per 1M) |
|---|---|---|---|---|
| Claude Sonnet 4 (Thinking) | 55.9% | 7.76% | Yes | $3 / $15 |
| Claude Opus 4 (Thinking) | 56.6% | 10.72% | Yes | $15 / $75 |
| Claude 3.7 Sonnet | — | 8.04% | Yes | $3 / $15 |
| o4-mini (high) | 80.2% | 18.08% | Yes | $1.10 / $4.40 |
Claude Sonnet 4 Thinking nearly matches Opus 4 Thinking on LiveCodeBench at $3/$15 — making it the better value for most coding tasks. o4-mini-high dominates on coding benchmarks at lower cost, making it the preferred choice for purely coding-intensive workloads. Claude Sonnet 4 Thinking is differentiated by Anthropic's safety properties and visible thinking.
Specs and scores from Anthropic's official Claude 4 announcement (May 2025) and Benchgen evaluations. Pricing cited to the Anthropic pricing page. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.