Benchgen

CharXiv Reasoning (No Tools) — Results

RankModelScore
1claude-fable-586.5
2kimi-k384.8
3gpt-5-6-sol84.7
4kimi-k2-680.4
5gemini-3-1-pro80.2
6inkling78.1
7kimi-k2-577.5

CharXiv Reasoning (No Tools)

1 phaseActive

CharXiv Reasoning evaluated without Python tool use — pure language model chart reasoning from arXiv figures. Metric: % correct. Differs from the 'with Python' variant on this platform.

Overview

CharXiv Reasoning (No Tools)

Category Metric Setting Saturation

Paper GitHub Dataset

Quick answer: CharXiv Reasoning (No Tools) is the CharXiv benchmark's reasoning sub-task evaluated without Python code execution — models must reason about scientific charts from arXiv papers using only language, without the ability to write and run Python for numerical computation. Scores are slightly lower than the "with Python" variant. Claude Fable 5 leads the Inkling comparison set at 86.5%.

At a Glance

What it tests: A model's ability to answer complex reasoning questions about real scientific charts extracted from arXiv papers, using only its language reasoning capabilities (no Python tool use allowed).

Why it matters: Many frontier models support Python code execution as a reasoning tool. Evaluating CharXiv Reasoning without Python isolates pure multimodal language reasoning ability from programming-assisted computation. This variant shows which models can reason about charts through understanding alone — a more fundamental capability.

Difference from the "with Python" variant: The CharXiv Reasoning (with Python) variant allows models to write and execute Python code to extract or compute values from charts. The "no tools" variant does not permit this, making it harder for questions requiring numerical computation.

Benchmark Specifications

FieldValue
Task categoryChart / scientific figure reasoning
Metric% accuracy
Tool useNone (no Python execution)
SourceCharXiv benchmark, Reasoning sub-task
Charts fromarXiv scientific papers
Created byZirui Wang, Mengzhou Xia, Luxi He, Howard Chen et al.
AffiliationPrinceton NLP
PaperCharXiv: Charting Gaps in Realistic Chart Understanding in Language Models (arXiv 2406.18521)
GitHubprinceton-nlp/CharXiv
Datasetprinceton-nlp/CharXiv on HuggingFace

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026). 6 models reported; Nemotron 3 Ultra, GLM 5.2, and DeepSeek V4 Pro not in this set.

RankModelScoreWeights
1Claude Fable 586.5%Closed
2GPT-5.6 Sol84.7%Closed
3Kimi K2.680.4%Open
4Gemini 3.1 Pro80.2%Closed
5Inkling78.1%Open
6Kimi K2.577.5%Open

Last updated 2026-07-16.