| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 63.3 |
1 phaseActive
Office document question-answering benchmark testing comprehension of Word, Excel, PowerPoint, and PDF files. Metric: % accuracy.
Quick answer: OfficeQA Pro tests a model's ability to answer questions grounded in real office documents — Word documents, Excel spreadsheets, PowerPoint presentations, and PDFs — requiring accurate extraction and reasoning over structured and semi-structured office content. Kimi K3 scores 63.3% as of July 2026.
What it tests: Question-answering accuracy over diverse office document formats, requiring both content extraction and reasoning across tables, slides, and formatted text.
Why it matters: Office documents are a dominant real-world data source for enterprise AI assistants. OfficeQA Pro targets this practical, high-value use case directly.
Known limitations: As an emerging benchmark, exact document sourcing and question design methodology are not yet independently published outside its citation by Moonshot AI.
OfficeQA Pro evaluates a model's ability to answer questions grounded in real office documents across common formats — Word, Excel, PowerPoint, and PDF. Questions require extracting and reasoning over structured elements (tables, cell references, slide content) as well as unstructured prose, testing practical document-understanding skills relevant to enterprise assistant use cases.
| Field | Value |
|---|---|
| Task category | Reasoning / document QA |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models answer questions about the content of provided office documents, scored on % accuracy against ground-truth answers.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 63.3% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run OfficeQA Pro.
| Benchmark | What it tests | Saturation |
|---|---|---|
| OfficeQA Pro | Office document question-answering | Low |
| SpreadsheetBench 2 | Spreadsheet-specific task completion | Low |
| OmniDocBench | Diverse PDF document parsing | Low |
| BIRD-SQL Dev | Structured data question-answering (SQL) | Medium |
Benchgen lets you run OfficeQA Pro against your own model, tracking document QA accuracy across Word, Excel, PowerPoint, and PDF formats.