| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 76.2 |
1 phaseActive
Rubric-graded benchmark evaluating the quality and completeness of AI-generated research reports. Metric: % rubric score.
Quick answer: ResearchRubrics evaluates AI-generated research reports against detailed expert-authored rubrics, scoring for completeness, accuracy, and structure rather than a single correct answer. Kimi K3 scores 76.2% as of July 2026.
What it tests: The quality and completeness of open-ended research report generation, judged against structured expert rubrics rather than a single ground-truth answer.
Why it matters: Deep-research tasks rarely have one correct output. Rubric-based grading captures nuanced quality dimensions (coverage, source use, structure) that simple accuracy metrics miss.
Known limitations: Rubric-based scoring can introduce grader variance depending on how strictly rubric criteria are applied, and exact rubric design is not independently published outside Kimi K3's own report.
ResearchRubrics evaluates a model's ability to produce comprehensive, well-structured research reports on open-ended topics. Each report is graded against an expert-authored rubric covering dimensions such as factual accuracy, source coverage, argument structure, and completeness, producing an aggregate rubric score rather than a binary correct/incorrect judgment.
| Field | Value |
|---|---|
| Task category | Agent / deep research |
| Metric | % rubric score |
| Saturation | Low |
| Created by | Not yet independently documented |
Generated research reports are graded against detailed expert rubrics covering multiple quality dimensions (accuracy, coverage, structure), with an aggregate % rubric score reflecting overall report quality.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 76.2% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run ResearchRubrics.
| Benchmark | What it tests | Saturation |
|---|---|---|
| ResearchRubrics | Rubric-graded research report quality | Low |
| DeepSearchQA | Multi-hop web research & synthesis | Low |
| BixBench | Biology research agent tasks | Low |
| AA-Briefcase | Professional knowledge-work quality | Low |
Benchgen lets you run ResearchRubrics-style evaluations against your own model, tracking rubric-graded research report quality over time.