Benchgen

TerminalBench 2.1 — Results

RankModelScore
1gpt-5-6-sol88.8
2kimi-k388.3
3qwen3-8-max86.6
4gpt-5-583.4
5muse-spark-1-282.9
6fugu-ultra82.1
7claude-sonnet-580.4
8fugu80.2
9claude-opus-4-771.7
10seed-2-1-pro71
11gemini-3-1-pro70.7
12laguna-s-2-170.2
13seed-2-1-turbo67.6
14claude-sonnet-4-667
15inkling63.8
16solar-pro-457
17muse-glimmer51.7
18kimi-k2-thinking-090547.1
19minimax-m246.3
20nova-2-pro41.3
21deepseek-v3-2-exp37.7
22glm-4-733.3
23deepseek-v3-131.3
24kimi-k2-instruct30
25nemotron-3-super-120b-a12b25.8
T

TerminalBench 2.1

1 phaseActive

Agentic benchmark testing AI models on long-horizon terminal and shell interactions — version 2.1 measures reliability across real CLI tasks.

Overview

TerminalBench 2.1

Category Metric Version Saturation

Quick answer: TerminalBench 2.1 is an agentic benchmark that evaluates AI models on long-horizon terminal and shell interaction tasks — measuring how reliably a model can navigate real command-line environments to complete multi-step goals. Fugu Ultra scores 82.1% and Fugu scores 80.2% on version 2.1 as of June 2026.

At a Glance

What it tests: An AI agent's ability to complete real, goal-oriented tasks through a terminal/shell interface — including file manipulation, process management, and multi-step CLI workflows.

Why it matters: Terminal proficiency is a critical building block for software engineering agents and autonomous developer tooling; TerminalBench provides a standardised measure of this capability.

Known limitations: Version-specific scoring means cross-version comparisons require care. Results are sensitive to the agent scaffold and tool access provided.

What TerminalBench 2.1 Measures

TerminalBench evaluates an agent's ability to autonomously complete structured tasks using terminal commands — including tasks that require sequential shell interactions, error recovery, and multi-step workflows. Version 2.1 represents a refresh of the task set and scoring criteria, maintaining the benchmark's focus on practical long-horizon terminal capability.

Scores are reported as a percentage of tasks successfully completed. A score above 80% indicates strong terminal-agent capability; frontier models typically cluster in the 70–82% range on version 2.1.

Benchmark Specifications

FieldValue
Task categoryAgent / terminal interaction
Metric% tasks completed successfully
Version2.1
SaturationLow

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu Ultra82.1%Sakana Fugu technical report2026-06
2Fugu80.2%Sakana Fugu technical report2026-06
3Fable 5 / Mythos Preview (max)74.6%Sakana Fugu technical report2026-06

Scores sourced from Sakana AI's Fugu technical report, June 2026.