| Rank | Model | Score |
|---|---|---|
| 1 | claude-opus-5-5 | 66.4 |
| 2 | claude-mythos-5-1 | 60.9 |
| 3 | gpt-6-astra | 57.9 |
| 4 | claude-fable-5-1 | 55.8 |
| 5 | grok-4-7 | 37.6 |
| 6 | deepseek-v4-1-flash | 31.2 |
1 phaseActive
Agentic terminal-use benchmark. Claude Mythos 5.1 leads at 60.9%, Claude Fable 5.1 scores 55.8% as of September 2026.
Quick answer: Terminal-Bench 4.0 is an agentic terminal/coding benchmark evaluating how well AI models complete multi-step tasks in a shell environment. Claude Mythos 5.1 leads at 60.9%, with Claude Fable 5.1 close behind at 55.8%, both ahead of Claude Opus 5 (52.3%) and GPT-5.6 Sol (37.3%), as reported by Anthropic in September 2026.
What it tests: An AI agent's ability to complete real, multi-step terminal/coding tasks — reading and modifying code, running shell commands, and verifying task completion — within a sandboxed environment.
Why it matters: Terminal Bench 4.0 is the version of the Terminal-Bench task suite referenced in Anthropic's September 2026 Claude Fable 5.1 / Claude Mythos 5.1 launch, providing a fresher, currently-unsaturated read on agentic coding/terminal capability than earlier Terminal-Bench releases.
Known limitations: Anthropic's announcement does not publicly detail the task set, authorship, or scoring methodology for this specific version; treat this page as tracking the scores reported in that release rather than an independently verified benchmark specification. It is a distinct, non-comparable benchmark from TerminalBench 2.1 and Terminal-Bench 3.0 despite the similar name.
Terminal-Bench 4.0 measures an agent's ability to autonomously complete goal-oriented tasks through a terminal/shell interface, in the same broad tradition as earlier Terminal-Bench releases — file manipulation, code changes, and multi-step CLI workflows, verified against the expected end state of the environment. Anthropic references it as one of the primary agentic-coding benchmarks in its Claude Fable 5.1 / Claude Mythos 5.1 launch, reporting scores for Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Opus 5, and GPT-5.6 Sol.
Notably, Claude Mythos 5.1 — the trusted-access-only, more-permissive-safeguard sibling of Fable 5.1 — reports the highest score of any evaluated model on this version of the benchmark, ahead of Fable 5.1 itself.
| Field | Value |
|---|---|
| Task category | Agent (terminal / coding) |
| Metric | % tasks completed successfully |
| Version | 4.0 |
| Saturation | Low |
Agents attempt terminal-based coding and shell tasks and are scored on the percentage of tasks successfully completed, verified via automated checks against the expected end state of the environment — consistent with the scoring approach used by earlier Terminal-Bench versions.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Claude Mythos 5.1 | 60.9% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 2 | Claude Fable 5.1 | 55.8% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 3 | Claude Opus 5 | 52.3% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 4 | Claude Fable 5 | 42.0% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 5 | GPT-5.6 Sol | 37.3% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
Scores sourced from Anthropic's Claude Fable 5.1 / Claude Mythos 5.1 announcement, September 2026.
No Benchgen results yet — be the first to run Terminal-Bench 4.0.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Terminal-Bench 4.0 | Agentic terminal/coding tasks (2026 revision) | Low |
| Terminal-Bench 3.0 | Harder, rolling agentic terminal-use tasks (Stanford/Laude Institute) | Low |
| TerminalBench 2.1 | Earlier-generation agentic terminal-use tasks | Medium |
| Terminal-Bench-Science 0.1 | Agentic scientific-research tasks via terminal use | Low |
Terminal-Bench 4.0, Terminal-Bench 3.0, and TerminalBench 2.1 are distinct, non-comparable benchmarks despite similar names — each uses its own task set and scoring scale.
Benchgen lets you run agentic terminal/coding tasks against your own model and harness, tracking task-completion rate over time as a complement to vendor-reported Terminal-Bench 4.0 scores.