Benchgen

Terminal-Bench 3.0 — Results

RankModelScore
1grok-4-626
T

Terminal-Bench 3.0

1 phaseActive

Stanford/Laude Institute's rolling, much harder successor to Terminal-Bench 2 — an ongoing agentic terminal-use benchmark, formerly Frontier-Bench.

Overview

Terminal-Bench 3.0

Category Metric Saturation Created

GitHub Leaderboard

Quick answer: Terminal-Bench 3.0 (formerly and still commonly called Frontier-Bench) is a rolling, continuously-updated agentic terminal-use benchmark from Stanford and the Laude Institute, run via the open-source Harbor framework. It is a substantially harder, ongoing successor to the earlier Terminal-Bench 2 / 2.1 releases — top models resolve well under half the tasks. Grok 4.6 scores 26% as of August 2026.

At a Glance

What it tests: An AI agent's ability to complete real terminal-based tasks — navigating a shell environment, running commands, and correctly completing multi-step system administration or engineering tasks.

Why it matters: As earlier Terminal-Bench versions (2.0, 2.1) approached saturation for frontier models, Terminal-Bench 3.0 introduces a harder, rolling task set designed to remain unsaturated as models improve, giving a clearer read on current frontier terminal-agent capability.

Known limitations: Because the dataset is a rolling, ongoing set (rather than a fixed static benchmark), scores from different dates are not always directly comparable — task composition can shift over time.

What Terminal-Bench 3.0 Measures

Terminal-Bench 3.0 is developed by Stanford researchers in partnership with the Laude Institute, and is run through Harbor, an open-source agent evaluation framework. It is the direct successor to Terminal-Bench 2.0 and 2.1 — a distinct, much harder benchmark with its own scale: while Terminal-Bench 2.1 tops out with frontier models resolving 70–80%+ of tasks, Terminal-Bench 3.0's top models resolve well under half the task set, confirming it uses a substantially harder and non-overlapping task distribution.

The project is also known informally as Frontier-Bench, and maintains a live, continuously-updated public leaderboard at frontierbench.ai.

Benchmark Specifications

FieldValue
Task categoryAgent (terminal use)
Metric% task resolution rate
SaturationLow
Created byStanford / Laude Institute
FrameworkHarbor
DatasetFrontier-Bench on Harbor Hub
Live leaderboardfrontierbench.ai

How Terminal-Bench 3.0 Is Scored

Agents attempt each terminal task inside a sandboxed shell environment and are scored on the percentage of tasks resolved correctly, verified via automated checks against the expected end state of the environment. Because the dataset is rolling, published scores are typically reported alongside a confidence interval and a date.

State-of-the-Art Results

RankModelScoreSourceDate
1Grok 4.626%xAI Grok 4.6 announcement2026-08

Score sourced from xAI's Grok 4.6 announcement, August 2026. See frontierbench.ai for the live, continuously-updated leaderboard across all evaluated model+harness combinations.

Terminal-Bench 3.0 on Benchgen

No Benchgen results yet — be the first to run Terminal-Bench 3.0.

Terminal-Bench 3.0 vs Other Benchmarks

BenchmarkWhat it testsSaturation
Terminal-Bench 3.0Harder, rolling agentic terminal-use tasksLow
TerminalBench 2.1Earlier-generation agentic terminal-use tasksMedium
OSWorld VerifiedComputer-use agent tasks in real OS environmentsLow

Terminal-Bench 3.0 and TerminalBench 2.1 are distinct, non-comparable benchmarks despite the similar name — 2.1 is closer to saturated for frontier models, while 3.0 is a deliberately harder, rolling replacement.

Run Terminal-Bench 3.0 on Your Model

Benchgen lets you run agentic terminal tasks against your own model and harness, tracking resolution rate over time as a complement to vendor-reported Terminal-Bench 3.0 scores.

Frequently Asked Questions

What is Terminal-Bench 3.0? Terminal-Bench 3.0 (also called Frontier-Bench) is a rolling agentic terminal-use benchmark from Stanford and the Laude Institute, run via the Harbor framework, and is a much harder successor to Terminal-Bench 2.1.
What does a good score look like on Terminal-Bench 3.0? Grok 4.6 reports 26% as of August 2026; top models on the public leaderboard generally resolve under half the task set, reflecting the benchmark's difficulty.
Who created Terminal-Bench 3.0? Terminal-Bench 3.0 was created by researchers at Stanford in partnership with the Laude Institute, using the open-source Harbor evaluation framework.
How does Terminal-Bench 3.0 differ from TerminalBench 2.1? They are distinct benchmarks despite the similar name: TerminalBench 2.1 is an earlier, easier task set closer to saturation, while Terminal-Bench 3.0 is a harder, ongoing rolling dataset where top models resolve well under half the tasks.