Benchgen

OCRBench v2 — Results

RankModelScore
1north-micro-vision-instruct36.7
O

OCRBench v2

1 phaseActive

Large-scale bilingual OCR benchmark with 4x more tasks than OCRBench v1 — text localization, handwriting, and logical reasoning over text-rich images.

Overview

OCRBench v2

Category Metric Tasks Saturation Created

Paper GitHub Leaderboard

Quick answer: OCRBench v2 (Fu et al., 2025) is a large-scale, bilingual (English/Chinese) text-centric benchmark with roughly 4x more tasks than the original OCRBench, covering 31 diverse scenarios and 10,000 human-verified question-answer pairs — including text localization, handwritten content extraction, and logical reasoning over text-rich images. North Micro Vision scores 36.7% on the English subset as of August 2026.

At a Glance

What it tests: Bilingual OCR capability across 31 scenarios and 10,000 human-verified questions, emphasizing harder sub-tasks than plain text recognition — text localization, handwriting extraction, fine-grained/layout perception, complex element parsing, and logical reasoning.

Why it matters: OCRBench v2's authors found most large multimodal models score below 50/100 on its harder task mix, showing that OCR capability claimed via simpler recognition benchmarks doesn't necessarily transfer to localization- and reasoning-heavy real-world OCR tasks.

Known limitations: As a newer, deliberately harder benchmark, current absolute scores are still low across the board — most models, including several frontier systems, score below 50%, so results should be read as early-stage progress rather than near-ceiling performance.

What OCRBench v2 Measures

OCRBench v2 substantially expands on the original OCRBench, offering roughly 4x more task types across 31 diverse real-world scenarios, with 10,000 human-verified question-answer pairs plus a held-out private test set of 1,500 manually annotated images used to validate result consistency. Task types go well beyond plain text recognition to include precise text localization (bounding a specific piece of text within an image), handwritten content extraction, fine-grained and layout perception, complex multi-element parsing, and logical reasoning over text-rich content.

The paper's central finding is that most large multimodal models — including many considered strong OCR performers on the original OCRBench — score below 50 out of 100 on this harder task mix, and share five common limitation categories: rare/uncommon text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning.

Because of this difficulty gap, OCRBench v2 is currently the more discriminative and future-proof OCR benchmark of the two, while the original OCRBench remains useful mainly as a simpler, faster baseline check.

Benchmark Specifications

FieldValue
Task categoryReasoning (OCR, text localization, layout)
Metric% accuracy (normalized; original scale 0-100)
Number of tasks10,000 (31 scenarios)
SaturationLow
Created byFu et al.
Source paperFu et al. 2025
GitHubYuliang-Liu/MultimodalOCR
Project page99franklin.github.io/ocrbench_v2

How OCRBench v2 Is Scored

Each of the 10,000 questions is scored for correctness across the benchmark's task categories (localization, extraction, layout perception, complex parsing, and reasoning); the overall score is the percent accuracy across the full public English or Chinese test subset, with a private held-out set used to validate result reliability.

State-of-the-Art Results

Score sourced from Cohere Labs' North Micro Vision Instruct launch announcement, August 2026, evaluated on the English (v2_en) subset.

OCRBench v2 on Benchgen

No Benchgen results yet — be the first to run OCRBench v2.

OCRBench v2 vs Other Benchmarks

BenchmarkWhat it testsSaturation
OCRBench v2Bilingual, 31-scenario OCR with localization and reasoningLow
OCRBenchFive-group OCR benchmark: recognition, scene/doc VQA, KIE, handwritingMedium
InfoVQADense infographics combining text, charts, and layoutMedium

OCRBench v2 is the harder, more comprehensive successor to OCRBench — use it when comparing frontier or near-frontier models where the original OCRBench no longer discriminates well.

Run OCRBench v2 on Your Model

Benchgen lets you evaluate your own vision-language model's OCR capability across text localization, handwriting extraction, and logical reasoning, and track results across model versions.

Frequently Asked Questions

What is OCRBench v2? OCRBench v2 is a large-scale, bilingual OCR benchmark spanning 31 scenarios and 10,000 questions, testing text localization, handwriting extraction, and logical reasoning — not just plain text recognition.
What does a good OCRBench v2 score look like? North Micro Vision Instruct reports 36.7% on the English subset as of August 2026; most large multimodal models, including several frontier systems, score below 50%.
Who created OCRBench v2? OCRBench v2 was created by Fu et al.; see the original paper.
Is OCRBench v2 saturated? No — it's a low-saturation, deliberately hard benchmark; most models today score well below 50 out of 100.