Benchgen

DocVQA — Results

RankModelScore
1deepseek-v4-1-flash95.6
2north-micro-vision-instruct92.1
3lfm2-5-vl-3b91.1
4claude-3-5-sonnet0.952
5qwen2-5-omni-7b0.952
6qwen2-5-vl-32b-instruct0.948
7llama-4-maverick0.944
8llama-4-scout0.944
9grok-20.936
10nova-pro0.935
11gpt-4o0.928
12nova-lite0.924
13llama-3-2-90b-instruct0.901
14gemma-3-12b0.871
15gemma-3-27b0.866
16grok-1-50.856
17gemma-3-4b0.758
D

DocVQA

1 phaseActive

50,000-question visual QA benchmark over 12,000+ real document images — tests document layout comprehension and information retrieval from scanned pages. Metric: ANLS.

Overview

DocVQA

Category Metric Saturation Tasks

Paper

Quick answer: DocVQA is a visual question answering benchmark by Mathew et al. (2020) containing 50,000 questions over 12,000+ real document images. It tests AI models on document layout comprehension, information retrieval, and understanding of tables, forms, and handwriting in scanned documents. Qwen2.5 VL 72B leads with 96.4% across 26 evaluated models.


What Does DocVQA Test?

DocVQA evaluates multimodal AI on realistic document understanding tasks. Documents include forms, invoices, reports, scientific articles, and other business documents with complex layouts. Models must understand both visual structure and textual content to answer questions correctly.

Document typeExamples
FormsTax forms, registration documents
InvoicesBusiness invoices, receipts
Scientific papersCharts, tables, figures
ReportsFinancial reports, company filings
Handwritten documentsPartially handwritten forms

How Is DocVQA Scored?

DocVQA uses ANLS (Average Normalized Levenshtein Similarity), which measures character-level similarity between the predicted and ground-truth answer. This handles OCR imperfections and slight answer variations more robustly than exact match. Scores are reported on a 0–1 scale.


Key Facts

PropertyValue
PublishedJuly 2020
Tasks50,000 questions
Images12,000+ documents
MetricANLS
Score range0–1
Top modelQwen2.5 VL 72B Instruct (0.964)
Models evaluated26

FAQ

What is DocVQA? DocVQA is a visual question answering benchmark for document images, containing 50,000 questions over 12,000+ documents including forms, invoices, and scientific papers. It tests AI models on document layout comprehension and information retrieval.

Who created DocVQA? DocVQA was created by Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar at CVIT, IIIT Hyderabad, published in July 2020 (arXiv 2007.00398).

What metric does DocVQA use? DocVQA uses ANLS (Average Normalized Levenshtein Similarity), which tolerates minor OCR errors and formatting differences more robustly than exact match scoring.

What score does the best model achieve on DocVQA? Qwen2.5 VL 72B Instruct currently leads with 0.964 (96.4%), followed by Qwen2.5 VL 7B Instruct at 0.957 and Claude 3.5 Sonnet at 0.952.