| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 44.2 |
1 phaseActive
Legal research benchmark testing case-law lookup, statutory interpretation, and citation accuracy. Metric: % accuracy.
Quick answer: Legal Research Bench tests AI models on legal research tasks — case-law lookup, statutory interpretation, and citation accuracy — evaluating whether a model can produce reliable, well-cited legal research output. Kimi K3 scores 44.2% as of July 2026.
What it tests: A model's ability to conduct accurate legal research, including correctly citing case law and statutes and interpreting legal text.
Why it matters: Legal research demands precise citation and interpretation; errors (including fabricated case citations) carry serious professional risk. This benchmark targets that specific reliability requirement.
Known limitations: As an emerging benchmark, exact task composition and jurisdiction coverage are not yet independently published outside its citation by Moonshot AI.
Legal Research Bench evaluates a model's ability to conduct legal research tasks — finding and citing relevant case law and statutes, interpreting legal language, and producing accurate legal analysis. It complements broader legal knowledge-work benchmarks (like Harvey Lab-AA) by focusing specifically on research and citation accuracy.
| Field | Value |
|---|---|
| Task category | Reasoning / legal domain |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models complete legal research tasks, scored on % accuracy against expert-verified reference answers, including correctness of legal citations and interpretations.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 44.2% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run Legal Research Bench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Legal Research Bench | Legal research & citation accuracy | Low |
| Harvey Lab-AA | Legal knowledge-work quality | Low |
| DeepSearchQA | Multi-hop web research & synthesis | Low |
| ResearchRubrics | Rubric-graded research report quality | Low |
Benchgen lets you run Legal Research Bench against your own model, tracking legal research accuracy over time before deploying to legal-domain use cases.