Benchgen

Legal Research Bench — Results

RankModelScore
1kimi-k344.2
L

Legal Research Bench

1 phaseActive

Legal research benchmark testing case-law lookup, statutory interpretation, and citation accuracy. Metric: % accuracy.

Overview

Category Metric Saturation Created

Quick answer: Legal Research Bench tests AI models on legal research tasks — case-law lookup, statutory interpretation, and citation accuracy — evaluating whether a model can produce reliable, well-cited legal research output. Kimi K3 scores 44.2% as of July 2026.

At a Glance

What it tests: A model's ability to conduct accurate legal research, including correctly citing case law and statutes and interpreting legal text.

Why it matters: Legal research demands precise citation and interpretation; errors (including fabricated case citations) carry serious professional risk. This benchmark targets that specific reliability requirement.

Known limitations: As an emerging benchmark, exact task composition and jurisdiction coverage are not yet independently published outside its citation by Moonshot AI.

Legal Research Bench evaluates a model's ability to conduct legal research tasks — finding and citing relevant case law and statutes, interpreting legal language, and producing accurate legal analysis. It complements broader legal knowledge-work benchmarks (like Harvey Lab-AA) by focusing specifically on research and citation accuracy.

Benchmark Specifications

FieldValue
Task categoryReasoning / legal domain
Metric% accuracy
SaturationLow
Created byNot yet independently documented

Models complete legal research tasks, scored on % accuracy against expert-verified reference answers, including correctness of legal citations and interpretations.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K344.2%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

No Benchgen results yet — be the first to run Legal Research Bench.

BenchmarkWhat it testsSaturation
Legal Research BenchLegal research & citation accuracyLow
Harvey Lab-AALegal knowledge-work qualityLow
DeepSearchQAMulti-hop web research & synthesisLow
ResearchRubricsRubric-graded research report qualityLow

Benchgen lets you run Legal Research Bench against your own model, tracking legal research accuracy over time before deploying to legal-domain use cases.

Frequently Asked Questions

What is Legal Research Bench? Legal Research Bench is a benchmark testing AI models on legal research tasks, including case-law lookup, statutory interpretation, and citation accuracy.
What does a good score look like on Legal Research Bench? Kimi K3 reports 44.2% as of July 2026, reflecting the significant remaining difficulty of fully reliable legal research for current frontier models.
Who created Legal Research Bench? Legal Research Bench's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.