Benchgen

ResearchRubrics — Results

RankModelScore
1kimi-k376.2
R

ResearchRubrics

1 phaseActive

Rubric-graded benchmark evaluating the quality and completeness of AI-generated research reports. Metric: % rubric score.

Overview

ResearchRubrics

Category Metric Saturation Created

Quick answer: ResearchRubrics evaluates AI-generated research reports against detailed expert-authored rubrics, scoring for completeness, accuracy, and structure rather than a single correct answer. Kimi K3 scores 76.2% as of July 2026.

At a Glance

What it tests: The quality and completeness of open-ended research report generation, judged against structured expert rubrics rather than a single ground-truth answer.

Why it matters: Deep-research tasks rarely have one correct output. Rubric-based grading captures nuanced quality dimensions (coverage, source use, structure) that simple accuracy metrics miss.

Known limitations: Rubric-based scoring can introduce grader variance depending on how strictly rubric criteria are applied, and exact rubric design is not independently published outside Kimi K3's own report.

What ResearchRubrics Measures

ResearchRubrics evaluates a model's ability to produce comprehensive, well-structured research reports on open-ended topics. Each report is graded against an expert-authored rubric covering dimensions such as factual accuracy, source coverage, argument structure, and completeness, producing an aggregate rubric score rather than a binary correct/incorrect judgment.

Benchmark Specifications

FieldValue
Task categoryAgent / deep research
Metric% rubric score
SaturationLow
Created byNot yet independently documented

How ResearchRubrics Is Scored

Generated research reports are graded against detailed expert rubrics covering multiple quality dimensions (accuracy, coverage, structure), with an aggregate % rubric score reflecting overall report quality.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K376.2%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

ResearchRubrics on Benchgen

No Benchgen results yet — be the first to run ResearchRubrics.

ResearchRubrics vs Other Benchmarks

BenchmarkWhat it testsSaturation
ResearchRubricsRubric-graded research report qualityLow
DeepSearchQAMulti-hop web research & synthesisLow
BixBenchBiology research agent tasksLow
AA-BriefcaseProfessional knowledge-work qualityLow

Run ResearchRubrics on Your Model

Benchgen lets you run ResearchRubrics-style evaluations against your own model, tracking rubric-graded research report quality over time.

Frequently Asked Questions

What is ResearchRubrics? ResearchRubrics is a benchmark that grades AI-generated research reports against expert-authored rubrics covering accuracy, coverage, and structure.
What does a good score look like on ResearchRubrics? Kimi K3 reports 76.2% as of July 2026, a strong result reflecting comprehensive, well-structured research report generation.
Who created ResearchRubrics? ResearchRubrics' originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.