| Rank | Model | Score |
|---|---|---|
| 1 | claude-opus-5-5 | 64.4 |
| 2 | claude-sonnet-5-5 | 61.6 |
| 3 | gpt-6-sol | 53.6 |
| 4 | claude-sonnet-5 | 15.6 |
1 phaseActive
Surge AI's visual chart recognition benchmark testing a model's ability to read and reason over charts directly, without code-execution tools.
Quick answer: Chartography is a visual chart recognition benchmark maintained by Surge AI that tests whether a model can read and reason over charts (bar, line, scatter, and other visualization types) directly from the image, without relying on code-execution or data-extraction tools. Claude Opus 5.5 leads the no-tools setting at 64.4% as of September 2026.
What it tests: A model's ability to extract values, trends, and relationships from a rendered chart image and answer questions about it, using vision alone rather than tool-assisted data extraction.
Why it matters: Chart reading is a common real-world knowledge-work task (financial reports, dashboards, scientific papers), and the no-tools setting isolates genuine visual reasoning from a model's ability to call out to code or OCR tools.
Known limitations: Chartography is run by a third party (Surge AI) rather than published as an open academic benchmark, so the task set and grading rubric are not independently inspectable; scores can also lag behind a model's latest version if the evaluator hasn't re-run it.
Chartography evaluates a model's ability to interpret data visualizations — reading axis values, identifying trends, comparing series, and answering quantitative questions about a chart — purely from the rendered image. Surge AI reports scores both "with tools" (where a model can use code execution or other aids) and "no tools" (vision-only), which can differ substantially for the same model and should not be compared across settings.
Because chart understanding sits at the intersection of vision and quantitative reasoning, labs increasingly cite Chartography alongside benchmarks like GDPval-AA and AA-Briefcase as evidence of real-world knowledge-work capability, rather than treating it as a narrow OCR test.
| Field | Value |
|---|---|
| Task category | Multimodal (visual chart reasoning) |
| Metric | % accuracy |
| Settings | No-tools (vision only) and with-tools (reported separately) |
| Saturation | Low |
| Created by | Surge AI |
| Dataset | Proprietary — not publicly released |
Models answer quantitative and qualitative questions about rendered charts, graded for accuracy against ground-truth values. The "no tools" setting (used for the primary leaderboard below) disables code execution and other extraction aids, isolating vision-based chart reading.
| Rank | Model | Score (no tools) | Source | Date |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 64.4% | Anthropic: Claude Sonnet 5.5 | 2026-09 |
| 2 | Claude Sonnet 5.5 | 61.6% | Anthropic: Claude Sonnet 5.5 | 2026-09 |
| 3 | GPT-6 Sol | 53.6% | Anthropic: Claude Sonnet 5.5 (via Surge AI) | 2026-09 |
| 4 | Claude Sonnet 5 | 15.6% | Anthropic: Claude Sonnet 5.5 | 2026-09 |
No-tools scores sourced from Anthropic's Claude Sonnet 5.5 announcement (September 2026), which also reports Claude Opus 5.5 and Claude Fable 5.1 "with tools" scores of 89.0% and 88.4% respectively — a separate, higher-scoring setting not directly comparable to the no-tools figures above. Anthropic notes the GPT-6 Sol figure (via Surge AI) may not yet reflect a post-launch image-understanding bugfix.
No Benchgen results yet — be the first to run Chartography.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Chartography | Visual chart reading and reasoning (no tools) | Low |
| GDPval-AA v2 | Real-world professional work across 44 occupations | Low |
| AA-Briefcase | Real-world knowledge-work document tasks | Low |
Benchgen lets you track multimodal chart-reasoning performance across model versions, complementing third-party-reported Chartography scores with independently repeatable evaluation.