Quick answer: Claude Sonnet 4 (released May 22, 2025) is Anthropic's mid-tier model from the Claude 4 generation, posting 72.7% on SWE-bench Verified — effectively matching the flagship Opus 4 — at $3/$15 per million tokens. A hybrid reasoning model with extended thinking, it was immediately adopted by GitHub Copilot as its primary coding agent backbone.
Where Claude Sonnet 4 leads
Where it lags
Best for: high-volume production coding agents and enterprise workflows needing Opus-level quality at sustainable cost.
Claude Sonnet 4 launched alongside Opus 4 on May 22, 2025. The headline story was the benchmark: 72.7% on SWE-bench Verified, barely behind Opus 4's 72.5% (the small difference is within measurement variance), at 5× lower cost. This made it the obvious default for teams building coding agents at scale, with Opus 4 reserved for workflows where reliability over very long runs justified the premium.
Anthropomorphically, it is described as a "significant upgrade from Sonnet 3.7" with enhanced steerability — meaning it follows system prompt constraints and complex multi-part instructions more faithfully, a critical property for agent systems where the model needs to stay on track across many steps. GitHub moved Copilot's coding agent to Sonnet 4 at launch.
| Field | Value |
|---|---|
| Organization | Anthropic |
| API identifier | claude-sonnet-4-20250514 |
| Context window | 200,000 tokens |
| Max output | 16,000 tokens |
| License | Proprietary |
| Release date | May 22, 2025 |
| Knowledge cutoff | March 2025 |
| Modality | Multimodal (text and vision) |
| Input (per 1M tokens) | Output (per 1M tokens) | |
|---|---|---|
| Anthropic | $3.00 | $15.00 |
Source: Anthropic pricing page.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| SWE-bench Verified | 72.7% | Anthropic — Introducing Claude 4 | 2025-05 |
| Terminal-bench 2.0 | 35.5% | Anthropic — Introducing Claude 4 | 2025-05 |
| GPQA Diamond (w/ extended thinking) | 76.0% | Anthropic — Introducing Claude 4 | 2025-05 |
Scores reported by Anthropic. Not Benchgen measurements.
| Model | SWE-bench Verified | GPQA Diamond | Price (in/out per 1M) |
|---|---|---|---|
| Claude Sonnet 4 | 72.7% | 76.0% | $3 / $15 |
| Claude Opus 4 | 72.5% | 79.1% | $15 / $75 |
| Claude Sonnet 4.6 | ~73%+ | — | $3 / $15 |
| GPT-5 | 74.9% | — | $1.25 / $10 |
Sonnet 4 delivers near-identical SWE-bench scores to Opus 4 at one-fifth the cost. Sonnet 4.6 improves on it at the same price. GPT-5 slightly leads on SWE-bench at lower output pricing. (Rival scores from respective announcements; not Benchgen measurements.)
Claude Sonnet 4 is a large language model developed by Anthropic.
This model isn’t on any benchmark leaderboard yet.