| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-6-sol | 88.8 |
| 2 | kimi-k3 | 88.3 |
| 3 | qwen3-8-max | 86.6 |
| 4 | gpt-5-5 | 83.4 |
| 5 | muse-spark-1-2 | 82.9 |
| 6 | fugu-ultra | 82.1 |
| 7 | claude-sonnet-5 | 80.4 |
| 8 | fugu | 80.2 |
| 9 | claude-opus-4-7 | 71.7 |
| 10 | seed-2-1-pro | 71 |
| 11 | gemini-3-1-pro | 70.7 |
| 12 | laguna-s-2-1 | 70.2 |
| 13 | seed-2-1-turbo | 67.6 |
| 14 | claude-sonnet-4-6 | 67 |
| 15 | inkling | 63.8 |
| 16 | solar-pro-4 | 57 |
| 17 | muse-glimmer | 51.7 |
| 18 | kimi-k2-thinking-0905 | 47.1 |
| 19 | minimax-m2 | 46.3 |
| 20 | nova-2-pro | 41.3 |
| 21 | deepseek-v3-2-exp | 37.7 |
| 22 | glm-4-7 | 33.3 |
| 23 | deepseek-v3-1 | 31.3 |
| 24 | kimi-k2-instruct | 30 |
| 25 | nemotron-3-super-120b-a12b | 25.8 |
1 phaseActive
Agentic benchmark testing AI models on long-horizon terminal and shell interactions — version 2.1 measures reliability across real CLI tasks.
Quick answer: TerminalBench 2.1 is an agentic benchmark that evaluates AI models on long-horizon terminal and shell interaction tasks — measuring how reliably a model can navigate real command-line environments to complete multi-step goals. Fugu Ultra scores 82.1% and Fugu scores 80.2% on version 2.1 as of June 2026.
What it tests: An AI agent's ability to complete real, goal-oriented tasks through a terminal/shell interface — including file manipulation, process management, and multi-step CLI workflows.
Why it matters: Terminal proficiency is a critical building block for software engineering agents and autonomous developer tooling; TerminalBench provides a standardised measure of this capability.
Known limitations: Version-specific scoring means cross-version comparisons require care. Results are sensitive to the agent scaffold and tool access provided.
TerminalBench evaluates an agent's ability to autonomously complete structured tasks using terminal commands — including tasks that require sequential shell interactions, error recovery, and multi-step workflows. Version 2.1 represents a refresh of the task set and scoring criteria, maintaining the benchmark's focus on practical long-horizon terminal capability.
Scores are reported as a percentage of tasks successfully completed. A score above 80% indicates strong terminal-agent capability; frontier models typically cluster in the 70–82% range on version 2.1.
| Field | Value |
|---|---|
| Task category | Agent / terminal interaction |
| Metric | % tasks completed successfully |
| Version | 2.1 |
| Saturation | Low |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu Ultra | 82.1% | Sakana Fugu technical report | 2026-06 |
| 2 | Fugu | 80.2% | Sakana Fugu technical report | 2026-06 |
| 3 | Fable 5 / Mythos Preview (max) | 74.6% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026.