Benchgen

Long Context Reasoning — Results

RankModelScore
1fugu74.7
2fugu-ultra73.3

Long Context Reasoning

1 phaseActive

Evaluates AI model reasoning across long input sequences — tests retention, multi-hop inference, and synthesis over extended contexts. Metric: % correct.

Overview

Long Context Reasoning

Category Metric Context Saturation

Quick answer: Long Context Reasoning evaluates an AI model's ability to perform multi-step reasoning across very long input sequences — testing whether models can retain relevant information, make multi-hop inferences, and synthesise answers from content spread across large documents or conversation histories. Fugu scores 74.7% and Fugu Ultra scores 73.3% as of June 2026.

At a Glance

What it tests: Reasoning quality over long contexts — including retrieval of relevant facts from distant positions, multi-hop inference chains, and synthesis of information across multiple sections of a long document.

Why it matters: Many enterprise AI use cases involve long documents, codebases, or conversation histories. A model that reasons accurately over long contexts is far more useful in production than one that degrades with length.

Known limitations: Performance is sensitive to context window size and attention architecture. Results may vary across different long-context task types (e.g., book QA vs. code reasoning vs. multi-document summarisation).

What Long Context Reasoning Measures

Long Context Reasoning benchmarks test a model's ability to maintain accurate reasoning over inputs that span tens of thousands to hundreds of thousands of tokens. Tasks typically require the model to retrieve specific facts from early in a long document, connect them with information elsewhere, and reason to a correct conclusion — rather than simply finding and repeating a nearby passage.

A score in the 70–75% range represents frontier performance as of mid-2026 and indicates strong long-context capability. Models that perform well here can be trusted to reason accurately over full codebases, legal documents, or extended research papers within their context window.

Benchmark Specifications

FieldValue
Task categoryReasoning over long contexts
Metric% correct
Context lengthLong (tens of thousands to 272K+ tokens)
SaturationLow

State-of-the-Art Results

RankModelScoreSourceDate
1Fugu74.7%Sakana Fugu technical report2026-06
2Fable 5 / Mythos Preview (max)74.3% (est.)Sakana Fugu technical report2026-06
3Fugu Ultra73.3%Sakana Fugu technical report2026-06

Scores sourced from Sakana AI's Fugu technical report, June 2026.