Benchgen

DocVQA — Results

RankModelScore
1north-micro-vision-instruct92.1
2lfm2-5-vl-3b91.1
3claude-3-5-sonnet0.952
4qwen2-5-omni-7b0.952
5qwen2-5-vl-32b-instruct0.948
6llama-4-maverick0.944
7llama-4-scout0.944
8grok-20.936
9nova-pro0.935
10gpt-4o0.928
11nova-lite0.924
12llama-3-2-90b-instruct0.901
13gemma-3-12b0.871
14gemma-3-27b0.866
15grok-1-50.856
16gemma-3-4b0.758
D

DocVQA

1 phaseActive

50,000-question visual QA benchmark over 12,000+ real document images — tests document layout comprehension and information retrieval from scanned pages. Metric: ANLS.

Overview

DocVQA

Category Metric Saturation Tasks

Paper

Quick answer: DocVQA is a visual question answering benchmark by Mathew et al. (2020) containing 50,000 questions over 12,000+ real document images. It tests AI models on document layout comprehension, information retrieval, and understanding of tables, forms, and handwriting in scanned documents. Qwen2.5 VL 72B leads with 96.4% across 26 evaluated models.


What Does DocVQA Test?

DocVQA evaluates multimodal AI on realistic document understanding tasks. Documents include forms, invoices, reports, scientific articles, and other business documents with complex layouts. Models must understand both visual structure and textual content to answer questions correctly.

Document typeExamples
FormsTax forms, registration documents
InvoicesBusiness invoices, receipts
Scientific papersCharts, tables, figures
ReportsFinancial reports, company filings
Handwritten documentsPartially handwritten forms

How Is DocVQA Scored?

DocVQA uses ANLS (Average Normalized Levenshtein Similarity), which measures character-level similarity between the predicted and ground-truth answer. This handles OCR imperfections and slight answer variations more robustly than exact match. Scores are reported on a 0–1 scale.


Key Facts

PropertyValue
PublishedJuly 2020
Tasks50,000 questions
Images12,000+ documents
MetricANLS
Score range0–1
Top modelQwen2.5 VL 72B Instruct (0.964)
Models evaluated26

FAQ

What is DocVQA? DocVQA is a visual question answering benchmark for document images, containing 50,000 questions over 12,000+ documents including forms, invoices, and scientific papers. It tests AI models on document layout comprehension and information retrieval.

Who created DocVQA? DocVQA was created by Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar at CVIT, IIIT Hyderabad, published in July 2020 (arXiv 2007.00398).

What metric does DocVQA use? DocVQA uses ANLS (Average Normalized Levenshtein Similarity), which tolerates minor OCR errors and formatting differences more robustly than exact match scoring.

What score does the best model achieve on DocVQA? Qwen2.5 VL 72B Instruct currently leads with 0.964 (96.4%), followed by Qwen2.5 VL 7B Instruct at 0.957 and Claude 3.5 Sonnet at 0.952.