| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 91.3 |
| 2 | fugu-ultra | 86.6 |
| 3 | fugu | 85.1 |
| 4 | inkling | 82 |
| 5 | muse-glimmer | 78.8 |
| 6 | kimi-k3 | 0.913 |
| 7 | claude-opus-4-7 | 0.91 |
| 8 | claude-opus-4-8 | 0.899 |
| 9 | muse-spark | 0.884 |
| 10 | claude-sonnet-5 | 0.883 |
| 11 | kimi-k2-6 | 0.867 |
| 12 | seed-2-1-pro | 0.864 |
| 13 | qwen3-7-plus | 0.859 |
| 14 | seed-2-1-turbo | 0.836 |
| 15 | gpt-5-2 | 0.821 |
| 16 | gpt-5-5 | 0.816 |
| 17 | qwen3-6-plus | 0.815 |
| 18 | gemini-3-pro | 0.814 |
| 19 | gpt-5 | 0.811 |
| 20 | gemini-3-flash | 0.803 |
| 21 | qwen3-5-27b | 0.795 |
| 22 | o3 | 0.786 |
| 23 | qwen3-6-27b | 0.784 |
| 24 | qwen3-6-35b-a3b | 0.78 |
| 25 | kimi-k2-5 | 0.775 |
1 phaseActive
Scientific chart reasoning benchmark from arXiv figures — tests AI models on complex visual data interpretation. Reasoning sub-task metric: % correct.
Quick answer: CharXiv Reasoning is the complex reasoning sub-task of the CharXiv benchmark (Wang et al., 2024), which evaluates AI models on scientific chart understanding using figures extracted from real arXiv papers. The Reasoning sub-task asks multi-step analytical questions that require integrating visual data with scientific reasoning. Fugu Ultra scores 86.6% and Fugu scores 85.1% as of June 2026.
What it tests: Complex, multi-step reasoning about scientific charts and figures drawn from real arXiv papers — including extracting quantitative data, comparing trends, and drawing analytical conclusions.
Why it matters: Real-world AI assistants for scientific research must interpret figures and charts, not just text. CharXiv Reasoning provides a rigorous, contamination-resistant measure of this capability using actual research figures.
Known limitations: Scientific figures from arXiv may be niche and domain-specific; performance likely varies by scientific domain. Requires vision capability — text-only models cannot be evaluated.
CharXiv extracts figures from recent arXiv preprints and generates two types of questions: descriptive (simple data reading) and reasoning (multi-step analytical). The Reasoning sub-task focuses on the harder analytical questions — for example, calculating percentage differences between trend lines, identifying crossover points, or drawing conclusions that require combining information from multiple chart elements.
By sourcing from recent arXiv preprints after model training cutoffs, CharXiv avoids contamination from training data. The reasoning sub-task is specifically challenging because it requires both accurate visual parsing and the ability to perform arithmetic or logical reasoning on extracted values.
| Field | Value |
|---|---|
| Task category | Reasoning / scientific figure understanding |
| Metric | % correct (accuracy) |
| Data source | Figures extracted from arXiv preprints |
| Task type | Complex multi-step analytical reasoning about charts |
| Saturation | Low |
| Created by | Wang et al. (Princeton NLP) |
| Source paper | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs (2024) |
| GitHub | princeton-nlp/CharXiv |
| Dataset | HuggingFace — princeton-nlp/CharXiv |
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Fugu Ultra | 86.6% | Sakana Fugu technical report | 2026-06 |
| 2 | Fugu | 85.1% | Sakana Fugu technical report | 2026-06 |
| 3 | Fable 5 / Mythos Preview (max) | 84.2% | Sakana Fugu technical report | 2026-06 |
Scores sourced from Sakana AI's Fugu technical report, June 2026.