Benchgen

Terminal-Bench-Science 0.1 — Results

RankModelScore
1gpt-6-astra64.6
2claude-opus-5-558.7
3claude-fable-5-152.6
T

Terminal-Bench-Science 0.1

1 phaseActive

Early-stage benchmark testing AI agents on agentic scientific-research tasks through a terminal interface. Metric: % accuracy, reported vs. cost.

Overview

Terminal-Bench-Science 0.1

Category Metric Version Saturation

Quick answer: Terminal-Bench-Science 0.1 is an early-version benchmark that measures how well AI agents perform agentic scientific-research tasks through a terminal interface, reporting accuracy against mean cost per task across multiple effort/reasoning levels. Claude Fable 5.1 scores 52.6% (at "max" effort) as of September 2026, per Anthropic's public reporting.

At a Glance

What it tests: An AI agent's ability to complete real scientific-research workflows — using a terminal/shell environment as the interaction surface — rather than isolated science-knowledge Q&A.

Why it matters: Most science benchmarks test static knowledge recall; Terminal-Bench-Science 0.1 instead measures whether an agent can actually execute a multi-step research task end-to-end, which is closer to how frontier labs expect models to accelerate real scientific work.

Known limitations: Version 0.1 signals this is an early-stage benchmark — task count and long-term stability are not yet publicly documented. Anthropic reports a standard error of ±3.5–4.5 points per model, and notes its own reproduction of published scores can differ slightly from a separate public leaderboard using a different harness (3 trials/task, Claude Code harness).

What Terminal-Bench-Science 0.1 Measures

Terminal-Bench-Science 0.1 evaluates an AI agent's ability to carry out agentic scientific-research tasks — the kind of multi-step, tool-using workflow a research assistant might perform — through a terminal/shell interface, rather than through static question answering. Results are typically reported as an accuracy-vs-cost curve across several reasoning-effort tiers (low, medium, high, xhigh, max), since higher-effort settings trade increased compute cost for higher accuracy.

Anthropic introduced this benchmark data point in its September 2026 Claude Fable 5.1 / Claude Mythos 5.1 announcement, reporting both its own "max effort" scores for Fable 5.1 and Fable 5, and noting that a separate public leaderboard (using 3 trials per task inside a Claude Code harness) reports different absolute numbers for Claude Opus 5 and Claude Fable 5 — within the benchmark's stated noise band, but a reminder that harness choice materially affects Terminal-Bench-Science 0.1 results.

Benchmark Specifications

FieldValue
Task categoryAgent (scientific research)
Metric% accuracy (reported vs. mean cost per task, USD)
Version0.1
SaturationLow
Reported standard error±3.5–4.5 pts per model

How Terminal-Bench-Science 0.1 Is Scored

Agents attempt scientific-research tasks inside a terminal environment and are scored on the percentage completed correctly. Because task difficulty allows models to spend more inference compute for higher accuracy, scores are commonly reported at multiple reasoning-effort levels (low/medium/high/xhigh/max) alongside the mean dollar cost per task at each level — making Terminal-Bench-Science 0.1 as much a cost-efficiency benchmark as a pure accuracy one.

State-of-the-Art Results

RankModelScoreSourceDate
1Claude Fable 5.152.6% (max effort)Anthropic: Claude Fable 5.1 and Mythos 5.12026-09
2Claude Opus 529.0% (Anthropic reproduction) / 30.0% (public leaderboard)Anthropic: Claude Fable 5.1 and Mythos 5.12026-09
3Claude Fable 524.7% (Anthropic reproduction, max) / 21.4% (public leaderboard)Anthropic: Claude Fable 5.1 and Mythos 5.12026-09
4GPT-5.6 Sol22.4%Anthropic: Claude Fable 5.1 and Mythos 5.12026-09

Scores sourced from Anthropic's Claude Fable 5.1 / Claude Mythos 5.1 announcement, September 2026. Fable 5.1's score is reported at "max" reasoning effort; lower-effort tiers trade accuracy for lower per-task cost (Fable 5.1: 26.3% at low effort up to 52.6% at max).

Terminal-Bench-Science 0.1 on Benchgen

No Benchgen results yet — be the first to run Terminal-Bench-Science 0.1.

Terminal-Bench-Science 0.1 vs Other Benchmarks

BenchmarkWhat it testsSaturation
Terminal-Bench-Science 0.1Agentic scientific-research tasks via terminal useLow
Terminal-Bench 4.0General agentic terminal/coding tasksLow
SciCodeScientific code-generation from research problemsLow
FrontierScience ResearchBroader frontier scientific-research task suiteLow

Run Terminal-Bench-Science 0.1 on Your Model

Benchgen lets you track agentic scientific-research performance across reasoning-effort tiers and harnesses, complementing vendor-reported Terminal-Bench-Science 0.1 scores with independently repeatable, cost-aware evaluation.

Frequently Asked Questions

What is Terminal-Bench-Science 0.1? Terminal-Bench-Science 0.1 is an early-version benchmark measuring how well AI agents complete agentic scientific-research tasks through a terminal interface, reported as accuracy versus mean cost per task.
What does a good score look like on Terminal-Bench-Science 0.1? Claude Fable 5.1 reports 52.6% at max reasoning effort as of September 2026 — the highest publicly reported score at benchmark launch, with a stated standard error of ±3.5–4.5 points per model.
Who created Terminal-Bench-Science 0.1? The benchmark's authorship isn't publicly detailed in Anthropic's launch announcement; Anthropic references a separate public leaderboard (3 trials/task, Claude Code harness) that also reports scores for this benchmark.