| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 64.6 |
| 2 | claude-opus-5-5 | 58.7 |
| 3 | claude-fable-5-1 | 52.6 |
1 phaseActive
Early-stage benchmark testing AI agents on agentic scientific-research tasks through a terminal interface. Metric: % accuracy, reported vs. cost.
Quick answer: Terminal-Bench-Science 0.1 is an early-version benchmark that measures how well AI agents perform agentic scientific-research tasks through a terminal interface, reporting accuracy against mean cost per task across multiple effort/reasoning levels. Claude Fable 5.1 scores 52.6% (at "max" effort) as of September 2026, per Anthropic's public reporting.
What it tests: An AI agent's ability to complete real scientific-research workflows — using a terminal/shell environment as the interaction surface — rather than isolated science-knowledge Q&A.
Why it matters: Most science benchmarks test static knowledge recall; Terminal-Bench-Science 0.1 instead measures whether an agent can actually execute a multi-step research task end-to-end, which is closer to how frontier labs expect models to accelerate real scientific work.
Known limitations: Version 0.1 signals this is an early-stage benchmark — task count and long-term stability are not yet publicly documented. Anthropic reports a standard error of ±3.5–4.5 points per model, and notes its own reproduction of published scores can differ slightly from a separate public leaderboard using a different harness (3 trials/task, Claude Code harness).
Terminal-Bench-Science 0.1 evaluates an AI agent's ability to carry out agentic scientific-research tasks — the kind of multi-step, tool-using workflow a research assistant might perform — through a terminal/shell interface, rather than through static question answering. Results are typically reported as an accuracy-vs-cost curve across several reasoning-effort tiers (low, medium, high, xhigh, max), since higher-effort settings trade increased compute cost for higher accuracy.
Anthropic introduced this benchmark data point in its September 2026 Claude Fable 5.1 / Claude Mythos 5.1 announcement, reporting both its own "max effort" scores for Fable 5.1 and Fable 5, and noting that a separate public leaderboard (using 3 trials per task inside a Claude Code harness) reports different absolute numbers for Claude Opus 5 and Claude Fable 5 — within the benchmark's stated noise band, but a reminder that harness choice materially affects Terminal-Bench-Science 0.1 results.
| Field | Value |
|---|---|
| Task category | Agent (scientific research) |
| Metric | % accuracy (reported vs. mean cost per task, USD) |
| Version | 0.1 |
| Saturation | Low |
| Reported standard error | ±3.5–4.5 pts per model |
Agents attempt scientific-research tasks inside a terminal environment and are scored on the percentage completed correctly. Because task difficulty allows models to spend more inference compute for higher accuracy, scores are commonly reported at multiple reasoning-effort levels (low/medium/high/xhigh/max) alongside the mean dollar cost per task at each level — making Terminal-Bench-Science 0.1 as much a cost-efficiency benchmark as a pure accuracy one.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 52.6% (max effort) | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 2 | Claude Opus 5 | 29.0% (Anthropic reproduction) / 30.0% (public leaderboard) | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 3 | Claude Fable 5 | 24.7% (Anthropic reproduction, max) / 21.4% (public leaderboard) | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
| 4 | GPT-5.6 Sol | 22.4% | Anthropic: Claude Fable 5.1 and Mythos 5.1 | 2026-09 |
Scores sourced from Anthropic's Claude Fable 5.1 / Claude Mythos 5.1 announcement, September 2026. Fable 5.1's score is reported at "max" reasoning effort; lower-effort tiers trade accuracy for lower per-task cost (Fable 5.1: 26.3% at low effort up to 52.6% at max).
No Benchgen results yet — be the first to run Terminal-Bench-Science 0.1.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Terminal-Bench-Science 0.1 | Agentic scientific-research tasks via terminal use | Low |
| Terminal-Bench 4.0 | General agentic terminal/coding tasks | Low |
| SciCode | Scientific code-generation from research problems | Low |
| FrontierScience Research | Broader frontier scientific-research task suite | Low |
Benchgen lets you track agentic scientific-research performance across reasoning-effort tiers and harnesses, complementing vendor-reported Terminal-Bench-Science 0.1 scores with independently repeatable, cost-aware evaluation.