| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 94.6 |
1 phaseActive
Legal knowledge-work benchmark testing AI models on legal research, drafting, and analysis quality. Metric: % accuracy.
Quick answer: Harvey Lab-AA is a legal knowledge-work benchmark evaluating AI models on tasks such as legal research, document analysis, and drafting quality — the type of work handled by legal AI assistants for law firms and in-house legal teams. Kimi K3 scores 94.6% as of July 2026.
What it tests: A model's ability to perform accurate legal research and analysis tasks, reflecting real legal-professional workflows.
Why it matters: Legal AI is a fast-growing, high-stakes application area where accuracy is critical. Harvey Lab-AA offers a domain-specific signal beyond general reasoning benchmarks.
Known limitations: As a benchmark referencing "Harvey" (a known legal AI company) and "AA" (suggesting an Artificial Analysis collaboration), exact methodology and task sourcing are not yet independently published outside its citation by Moonshot AI.
Harvey Lab-AA evaluates a model's performance on legal knowledge-work tasks — such as case research, contract analysis, and legal document drafting — testing domain-specific accuracy and reasoning quality relevant to legal professional use cases.
| Field | Value |
|---|---|
| Task category | Reasoning / legal domain |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Not yet independently documented |
Models complete legal research and analysis tasks, scored on % accuracy against expert-verified reference answers or evaluation criteria.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 94.6% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run Harvey Lab-AA.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Harvey Lab-AA | Legal knowledge-work quality | Low |
| Legal Research Bench | Legal research task accuracy | Low |
| AA-Briefcase | Professional knowledge-work quality | Low |
| CorpFin v2 | Corporate finance task automation | Low |
Benchgen lets you run Harvey Lab-AA against your own model, tracking legal knowledge-work accuracy over time as you evaluate deployment for legal-domain use cases.