| Rank | Model | Score |
|---|---|---|
| 1 | claude-fable-5 | 86.5 |
| 2 | kimi-k3 | 84.8 |
| 3 | gpt-5-6-sol | 84.7 |
| 4 | kimi-k2-6 | 80.4 |
| 5 | gemini-3-1-pro | 80.2 |
| 6 | inkling | 78.1 |
| 7 | kimi-k2-5 | 77.5 |
1 phaseActive
CharXiv Reasoning evaluated without Python tool use — pure language model chart reasoning from arXiv figures. Metric: % correct. Differs from the 'with Python' variant on this platform.
Quick answer: CharXiv Reasoning (No Tools) is the CharXiv benchmark's reasoning sub-task evaluated without Python code execution — models must reason about scientific charts from arXiv papers using only language, without the ability to write and run Python for numerical computation. Scores are slightly lower than the "with Python" variant. Claude Fable 5 leads the Inkling comparison set at 86.5%.
What it tests: A model's ability to answer complex reasoning questions about real scientific charts extracted from arXiv papers, using only its language reasoning capabilities (no Python tool use allowed).
Why it matters: Many frontier models support Python code execution as a reasoning tool. Evaluating CharXiv Reasoning without Python isolates pure multimodal language reasoning ability from programming-assisted computation. This variant shows which models can reason about charts through understanding alone — a more fundamental capability.
Difference from the "with Python" variant: The CharXiv Reasoning (with Python) variant allows models to write and execute Python code to extract or compute values from charts. The "no tools" variant does not permit this, making it harder for questions requiring numerical computation.
| Field | Value |
|---|---|
| Task category | Chart / scientific figure reasoning |
| Metric | % accuracy |
| Tool use | None (no Python execution) |
| Source | CharXiv benchmark, Reasoning sub-task |
| Charts from | arXiv scientific papers |
| Created by | Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen et al. |
| Affiliation | Princeton NLP |
| Paper | CharXiv: Charting Gaps in Realistic Chart Understanding in Language Models (arXiv 2406.18521) |
| GitHub | princeton-nlp/CharXiv |
| Dataset | princeton-nlp/CharXiv on HuggingFace |
Scores from Inkling model card (Thinking Machines Lab, July 2026). 6 models reported; Nemotron 3 Ultra, GLM 5.2, and DeepSeek V4 Pro not in this set.
| Rank | Model | Score | Weights |
|---|---|---|---|
| 1 | Claude Fable 5 | 86.5% | Closed |
| 2 | GPT-5.6 Sol | 84.7% | Closed |
| 3 | Kimi K2.6 | 80.4% | Open |
| 4 | Gemini 3.1 Pro | 80.2% | Closed |
| 5 | Inkling | 78.1% | Open |
| 6 | Kimi K2.5 | 77.5% | Open |
Last updated 2026-07-16.