Benchgen
Models/anthropic/

Claude Sonnet 4

DraftPublic

Model Details

Claude Sonnet 4

Organization Context Pricing License Modality Released

Quick answer: Claude Sonnet 4 (released May 22, 2025) is Anthropic's mid-tier model from the Claude 4 generation, posting 72.7% on SWE-bench Verified — effectively matching the flagship Opus 4 — at $3/$15 per million tokens. A hybrid reasoning model with extended thinking, it was immediately adopted by GitHub Copilot as its primary coding agent backbone.

At a Glance

Where Claude Sonnet 4 leads

  • 72.7% SWE-bench Verified at $3/$15 — delivers Opus-class coding performance at Sonnet pricing
  • GitHub Copilot's choice for its new coding agent, with Manus and iGent reporting significant gains over Sonnet 3.7
  • Enhanced steerability — follows complex instructions more precisely with fewer navigation errors
  • Extended thinking with tool use for tasks requiring deep reasoning

Where it lags

  • Superseded by Sonnet 4.5 and Sonnet 4.6 for new builds
  • 200K context (Sonnet 4.6 offers 1M in beta)

Best for: high-volume production coding agents and enterprise workflows needing Opus-level quality at sustainable cost.

What Claude Sonnet 4 Is

Claude Sonnet 4 launched alongside Opus 4 on May 22, 2025. The headline story was the benchmark: 72.7% on SWE-bench Verified, barely behind Opus 4's 72.5% (the small difference is within measurement variance), at 5× lower cost. This made it the obvious default for teams building coding agents at scale, with Opus 4 reserved for workflows where reliability over very long runs justified the premium.

Anthropomorphically, it is described as a "significant upgrade from Sonnet 3.7" with enhanced steerability — meaning it follows system prompt constraints and complex multi-part instructions more faithfully, a critical property for agent systems where the model needs to stay on track across many steps. GitHub moved Copilot's coding agent to Sonnet 4 at launch.

Specifications

FieldValue
OrganizationAnthropic
API identifierclaude-sonnet-4-20250514
Context window200,000 tokens
Max output16,000 tokens
LicenseProprietary
Release dateMay 22, 2025
Knowledge cutoffMarch 2025
ModalityMultimodal (text and vision)

Pricing

Input (per 1M tokens)Output (per 1M tokens)
Anthropic$3.00$15.00

Source: Anthropic pricing page.

Public Benchmark Scores

BenchmarkScoreSourceDate
SWE-bench Verified72.7%Anthropic — Introducing Claude 42025-05
Terminal-bench 2.035.5%Anthropic — Introducing Claude 42025-05
GPQA Diamond (w/ extended thinking)76.0%Anthropic — Introducing Claude 42025-05

Scores reported by Anthropic. Not Benchgen measurements.

Claude Sonnet 4 vs Alternatives

ModelSWE-bench VerifiedGPQA DiamondPrice (in/out per 1M)
Claude Sonnet 472.7%76.0%$3 / $15
Claude Opus 472.5%79.1%$15 / $75
Claude Sonnet 4.6~73%+$3 / $15
GPT-574.9%$1.25 / $10

Sonnet 4 delivers near-identical SWE-bench scores to Opus 4 at one-fifth the cost. Sonnet 4.6 improves on it at the same price. GPT-5 slightly leads on SWE-bench at lower output pricing. (Rival scores from respective announcements; not Benchgen measurements.)

Frequently Asked Questions

What is Claude Sonnet 4? Claude Sonnet 4 is Anthropic's mid-tier model from the Claude 4 generation, released May 22, 2025. It matches Claude Opus 4's SWE-bench Verified score (72.7% vs 72.5%) at 5× lower cost, making it the default choice for production coding agents.
How much does Claude Sonnet 4 cost? $3 per million input tokens and $15 per million output tokens.

Specs sourced from Anthropic's Claude 4 announcement. Last updated 2026-06-19.

Claude Sonnet 4

Claude Sonnet 4 is a large language model developed by Anthropic.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.