| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 94.3 |
1 phaseActive
Visual mathematical reasoning benchmark with 3,040 competition-sourced problems across 16 topics and 5 difficulty levels. Metric: % accuracy (pass@1).
Quick answer: MathVision (Wang et al., 2024) is a curated collection of 3,040 visually-contexted mathematical problems sourced from real math competitions, spanning 16 distinct mathematical topics and graded across 5 difficulty levels. Unlike text-only math benchmarks, every problem requires interpreting a diagram, chart, or geometric figure alongside the text. Kimi K3 scores 94.3% (pass@1) as of July 2026.
What it tests: A model's ability to combine visual perception (reading diagrams, graphs, geometric figures) with rigorous mathematical reasoning to solve competition-level problems.
Why it matters: Text-only math benchmarks like MATH and GSM8K don't test whether a model can extract the right information from a figure. MathVision fills this gap and revealed a substantial gap between top closed-source and open-source models at release.
Known limitations: Being sourced from real competitions, some problems may have partial overlap with models' pretraining data, though the visual component reduces straightforward memorization risk.
MathVision meticulously curates 3,040 high-quality mathematical problems with visual contexts from real math competitions, covering 16 mathematical topics (e.g., algebra, geometry, combinatorics) across 5 levels of difficulty. This establishes a comprehensive and diverse set of visual math challenges, enabling a more rigorous evaluation of models' true mathematical reasoning abilities in multimodal settings, rather than reasoning derived from purely textual problem statements.
| Field | Value |
|---|---|
| Task category | Math / visual reasoning |
| Metric | % accuracy (pass@1) |
| Number of problems | 3,040 |
| Topics | 16 |
| Difficulty levels | 5 |
| Saturation | Low |
| Created by | Wang et al. |
| Source paper | MathVision (arXiv 2402.14804) |
Models solve each problem given its accompanying image and text, and are scored on % accuracy (pass@1) against the ground-truth answer. Some evaluations also report pass@5 or majority-vote scores to characterize consistency across multiple sampled attempts.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 94.3% (pass@1) | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run MathVision.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| MathVision | Visual mathematical reasoning | 3,040 | Low |
| MATH | Text-only competition math | — | Low |
| AIME 2026 | Competition math (text) | — | Low |
| ZeroBench | Extremely hard visual reasoning | 100 | Low |
Benchgen lets you run MathVision against your own multimodal model, tracking visual math reasoning accuracy across topics and difficulty levels to guide targeted fine-tuning.