| Rank | Model | Score |
|---|---|---|
| 1 | grok-4-6 | 26 |
1 phaseActive
Stanford/Laude Institute's rolling, much harder successor to Terminal-Bench 2 — an ongoing agentic terminal-use benchmark, formerly Frontier-Bench.
Quick answer: Terminal-Bench 3.0 (formerly and still commonly called Frontier-Bench) is a rolling, continuously-updated agentic terminal-use benchmark from Stanford and the Laude Institute, run via the open-source Harbor framework. It is a substantially harder, ongoing successor to the earlier Terminal-Bench 2 / 2.1 releases — top models resolve well under half the tasks. Grok 4.6 scores 26% as of August 2026.
What it tests: An AI agent's ability to complete real terminal-based tasks — navigating a shell environment, running commands, and correctly completing multi-step system administration or engineering tasks.
Why it matters: As earlier Terminal-Bench versions (2.0, 2.1) approached saturation for frontier models, Terminal-Bench 3.0 introduces a harder, rolling task set designed to remain unsaturated as models improve, giving a clearer read on current frontier terminal-agent capability.
Known limitations: Because the dataset is a rolling, ongoing set (rather than a fixed static benchmark), scores from different dates are not always directly comparable — task composition can shift over time.
Terminal-Bench 3.0 is developed by Stanford researchers in partnership with the Laude Institute, and is run through Harbor, an open-source agent evaluation framework. It is the direct successor to Terminal-Bench 2.0 and 2.1 — a distinct, much harder benchmark with its own scale: while Terminal-Bench 2.1 tops out with frontier models resolving 70–80%+ of tasks, Terminal-Bench 3.0's top models resolve well under half the task set, confirming it uses a substantially harder and non-overlapping task distribution.
The project is also known informally as Frontier-Bench, and maintains a live, continuously-updated public leaderboard at frontierbench.ai.
| Field | Value |
|---|---|
| Task category | Agent (terminal use) |
| Metric | % task resolution rate |
| Saturation | Low |
| Created by | Stanford / Laude Institute |
| Framework | Harbor |
| Dataset | Frontier-Bench on Harbor Hub |
| Live leaderboard | frontierbench.ai |
Agents attempt each terminal task inside a sandboxed shell environment and are scored on the percentage of tasks resolved correctly, verified via automated checks against the expected end state of the environment. Because the dataset is rolling, published scores are typically reported alongside a confidence interval and a date.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Grok 4.6 | 26% | xAI Grok 4.6 announcement | 2026-08 |
Score sourced from xAI's Grok 4.6 announcement, August 2026. See frontierbench.ai for the live, continuously-updated leaderboard across all evaluated model+harness combinations.
No Benchgen results yet — be the first to run Terminal-Bench 3.0.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Terminal-Bench 3.0 | Harder, rolling agentic terminal-use tasks | Low |
| TerminalBench 2.1 | Earlier-generation agentic terminal-use tasks | Medium |
| OSWorld Verified | Computer-use agent tasks in real OS environments | Low |
Terminal-Bench 3.0 and TerminalBench 2.1 are distinct, non-comparable benchmarks despite the similar name — 2.1 is closer to saturated for frontier models, while 3.0 is a deliberately harder, rolling replacement.
Benchgen lets you run agentic terminal tasks against your own model and harness, tracking resolution rate over time as a complement to vendor-reported Terminal-Bench 3.0 scores.