| Rank | Model | Score |
|---|---|---|
| 1 | fugu | 74.7 |
| 2 | fugu-ultra | 73.3 |
1 phaseActive
Evaluates AI model reasoning across long input sequences — tests retention, multi-hop inference, and synthesis over extended contexts. Metric: % correct.
Quick answer: Long Context Reasoning evaluates an AI model's ability to perform multi-step reasoning across very long input sequences — testing whether models can retain relevant information, make multi-hop inferences, and synthesise answers from content spread across large documents or conversation histories. Fugu scores 74.7% and Fugu Ultra scores 73.3% as of June 2026.
What it tests: Reasoning quality over long contexts — including retrieval of relevant facts from distant positions, multi-hop inference chains, and synthesis of information across multiple sections of a long document.
Why it matters: Many enterprise AI use cases involve long documents, codebases, or conversation histories. A model that reasons accurately over long contexts is far more useful in production than one that degrades with length.
Known limitations: Performance is sensitive to context window size and attention architecture. Results may vary across different long-context task types (e.g., book QA vs. code reasoning vs. multi-document summarisation).
Long Context Reasoning benchmarks test a model's ability to maintain accurate reasoning over inputs that span tens of thousands to hundreds of thousands of tokens. Tasks typically require the model to retrieve specific facts from early in a long document, connect them with information elsewhere, and reason to a correct conclusion — rather than simply finding and repeating a nearby passage.
A score in the 70–75% range represents frontier performance as of mid-2026 and indicates strong long-context capability. Models that perform well here can be trusted to reason accurately over full codebases, legal documents, or extended research papers within their context window.
| Field | Value |
|---|---|
| Task category | Reasoning over long contexts |
| Metric | % correct |
| Context length | Long (tens of thousands to 272K+ tokens) |
| Saturation | Low |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu | 74.7% | Sakana Fugu technical report | 2026-06 |
| 2 | Fable 5 / Mythos Preview (max) | 74.3% (est.) | Sakana Fugu technical report | 2026-06 |
| 3 | Fugu Ultra | 73.3% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026.