| Rank | Model | Score |
|---|---|---|
| 1 | deepseek-v4-1-flash | 90.6 |
| 2 | deepseek-v4-1-flash | 90.6 |
| 3 | mimo-v2-6-pro-rl | 89.9 |
| 4 | gpt-5-6-sol | 88.8 |
| 5 | kimi-k3 | 88.3 |
| 6 | glm-5-3 | 88.2 |
| 7 | deepseek-v4-pro | 87.9 |
| 8 | mimo-v2-6-flash-rl | 87.6 |
| 9 | qwen3-8-max | 86.6 |
| 10 | ornith-1-5-397b | 86.1 |
| 11 | nex-n2-5-max | 86.1 |
| 12 | hy4-preview | 85.4 |
| 13 | glm-5-3-flash | 84.3 |
| 14 | deepseek-v4-flash-vision-exp | 83.9 |
| 15 | gpt-5-5 | 83.4 |
| 16 | muse-spark-1-2 | 82.9 |
| 17 | nex-n2-5-pro | 82.7 |
| 18 | fugu-ultra | 82.1 |
| 19 | claude-sonnet-5 | 80.4 |
| 20 | fugu | 80.2 |
| 21 | dots3-note-preview | 75.1 |
| 22 | qwen3-8-27b | 73 |
| 23 | claude-opus-4-7 | 71.7 |
| 24 | seed-2-1-pro | 71 |
| 25 | apodex-1-1 | 70.8 |
1 phaseActive
Agentic benchmark testing AI models on long-horizon terminal and shell interactions — version 2.1 measures reliability across real CLI tasks.
Quick answer: TerminalBench 2.1 is an agentic benchmark that evaluates AI models on long-horizon terminal and shell interaction tasks — measuring how reliably a model can navigate real command-line environments to complete multi-step goals. Fugu Ultra scores 82.1% and Fugu scores 80.2% on version 2.1 as of June 2026.
What it tests: An AI agent's ability to complete real, goal-oriented tasks through a terminal/shell interface — including file manipulation, process management, and multi-step CLI workflows.
Why it matters: Terminal proficiency is a critical building block for software engineering agents and autonomous developer tooling; TerminalBench provides a standardised measure of this capability.
Known limitations: Version-specific scoring means cross-version comparisons require care. Results are sensitive to the agent scaffold and tool access provided.
TerminalBench evaluates an agent's ability to autonomously complete structured tasks using terminal commands — including tasks that require sequential shell interactions, error recovery, and multi-step workflows. Version 2.1 represents a refresh of the task set and scoring criteria, maintaining the benchmark's focus on practical long-horizon terminal capability.
Scores are reported as a percentage of tasks successfully completed. A score above 80% indicates strong terminal-agent capability; frontier models typically cluster in the 70–82% range on version 2.1.
| Field | Value |
|---|---|
| Task category | Agent / terminal interaction |
| Metric | % tasks completed successfully |
| Version | 2.1 |
| Saturation | Low |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu Ultra | 82.1% | Sakana Fugu technical report | 2026-06 |
| 2 | Fugu | 80.2% | Sakana Fugu technical report | 2026-06 |
| 3 | Fable 5 / Mythos Preview (max) | 74.6% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026.