Benchgen

Terminal-Bench 4.0 — Results

RankModelScore
1claude-opus-5-566.4
2claude-mythos-5-160.9
3gpt-6-astra57.9
4claude-fable-5-155.8
5grok-4-737.6
6deepseek-v4-1-flash31.2
T

Terminal-Bench 4.0

1 phaseActive

Agentic terminal-use benchmark. Claude Mythos 5.1 leads at 60.9%, Claude Fable 5.1 scores 55.8% as of September 2026.

Overview

Terminal-Bench 4.0

Category Metric Version Saturation

Quick answer: Terminal-Bench 4.0 is an agentic terminal/coding benchmark evaluating how well AI models complete multi-step tasks in a shell environment. Claude Mythos 5.1 leads at 60.9%, with Claude Fable 5.1 close behind at 55.8%, both ahead of Claude Opus 5 (52.3%) and GPT-5.6 Sol (37.3%), as reported by Anthropic in September 2026.

At a Glance

What it tests: An AI agent's ability to complete real, multi-step terminal/coding tasks — reading and modifying code, running shell commands, and verifying task completion — within a sandboxed environment.

Why it matters: Terminal Bench 4.0 is the version of the Terminal-Bench task suite referenced in Anthropic's September 2026 Claude Fable 5.1 / Claude Mythos 5.1 launch, providing a fresher, currently-unsaturated read on agentic coding/terminal capability than earlier Terminal-Bench releases.

Known limitations: Anthropic's announcement does not publicly detail the task set, authorship, or scoring methodology for this specific version; treat this page as tracking the scores reported in that release rather than an independently verified benchmark specification. It is a distinct, non-comparable benchmark from TerminalBench 2.1 and Terminal-Bench 3.0 despite the similar name.

What Terminal-Bench 4.0 Measures

Terminal-Bench 4.0 measures an agent's ability to autonomously complete goal-oriented tasks through a terminal/shell interface, in the same broad tradition as earlier Terminal-Bench releases — file manipulation, code changes, and multi-step CLI workflows, verified against the expected end state of the environment. Anthropic references it as one of the primary agentic-coding benchmarks in its Claude Fable 5.1 / Claude Mythos 5.1 launch, reporting scores for Claude Fable 5.1, Claude Mythos 5.1, Claude Fable 5, Claude Opus 5, and GPT-5.6 Sol.

Notably, Claude Mythos 5.1 — the trusted-access-only, more-permissive-safeguard sibling of Fable 5.1 — reports the highest score of any evaluated model on this version of the benchmark, ahead of Fable 5.1 itself.

Benchmark Specifications

FieldValue
Task categoryAgent (terminal / coding)
Metric% tasks completed successfully
Version4.0
SaturationLow

How Terminal-Bench 4.0 Is Scored

Agents attempt terminal-based coding and shell tasks and are scored on the percentage of tasks successfully completed, verified via automated checks against the expected end state of the environment — consistent with the scoring approach used by earlier Terminal-Bench versions.

State-of-the-Art Results

Scores sourced from Anthropic's Claude Fable 5.1 / Claude Mythos 5.1 announcement, September 2026.

Terminal-Bench 4.0 on Benchgen

No Benchgen results yet — be the first to run Terminal-Bench 4.0.

Terminal-Bench 4.0 vs Other Benchmarks

BenchmarkWhat it testsSaturation
Terminal-Bench 4.0Agentic terminal/coding tasks (2026 revision)Low
Terminal-Bench 3.0Harder, rolling agentic terminal-use tasks (Stanford/Laude Institute)Low
TerminalBench 2.1Earlier-generation agentic terminal-use tasksMedium
Terminal-Bench-Science 0.1Agentic scientific-research tasks via terminal useLow

Terminal-Bench 4.0, Terminal-Bench 3.0, and TerminalBench 2.1 are distinct, non-comparable benchmarks despite similar names — each uses its own task set and scoring scale.

Run Terminal-Bench 4.0 on Your Model

Benchgen lets you run agentic terminal/coding tasks against your own model and harness, tracking task-completion rate over time as a complement to vendor-reported Terminal-Bench 4.0 scores.

Frequently Asked Questions

What is Terminal-Bench 4.0? Terminal-Bench 4.0 is an agentic terminal/coding benchmark referenced in Anthropic's September 2026 Claude Fable 5.1 / Claude Mythos 5.1 launch, measuring how reliably a model completes multi-step tasks in a shell environment.
What does a good score look like on Terminal-Bench 4.0? Claude Mythos 5.1 reports the highest score at 60.9%, followed by Claude Fable 5.1 at 55.8%, Claude Opus 5 at 52.3%, Claude Fable 5 at 42.0%, and GPT-5.6 Sol at 37.3%, all as of September 2026.
How does Terminal-Bench 4.0 differ from Terminal-Bench 3.0 and TerminalBench 2.1? They are distinct, non-comparable benchmarks despite the similar naming — each version uses its own task set and difficulty scale, so scores should not be compared across versions.