Benchgen

Harvey Lab-AA — Results

RankModelScore
1kimi-k394.6
H

Harvey Lab-AA

1 phaseActive

Legal knowledge-work benchmark testing AI models on legal research, drafting, and analysis quality. Metric: % accuracy.

Overview

Harvey Lab-AA

Category Metric Saturation Created

Quick answer: Harvey Lab-AA is a legal knowledge-work benchmark evaluating AI models on tasks such as legal research, document analysis, and drafting quality — the type of work handled by legal AI assistants for law firms and in-house legal teams. Kimi K3 scores 94.6% as of July 2026.

At a Glance

What it tests: A model's ability to perform accurate legal research and analysis tasks, reflecting real legal-professional workflows.

Why it matters: Legal AI is a fast-growing, high-stakes application area where accuracy is critical. Harvey Lab-AA offers a domain-specific signal beyond general reasoning benchmarks.

Known limitations: As a benchmark referencing "Harvey" (a known legal AI company) and "AA" (suggesting an Artificial Analysis collaboration), exact methodology and task sourcing are not yet independently published outside its citation by Moonshot AI.

What Harvey Lab-AA Measures

Harvey Lab-AA evaluates a model's performance on legal knowledge-work tasks — such as case research, contract analysis, and legal document drafting — testing domain-specific accuracy and reasoning quality relevant to legal professional use cases.

Benchmark Specifications

FieldValue
Task categoryReasoning / legal domain
Metric% accuracy
SaturationLow
Created byNot yet independently documented

How Harvey Lab-AA Is Scored

Models complete legal research and analysis tasks, scored on % accuracy against expert-verified reference answers or evaluation criteria.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K394.6%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

Harvey Lab-AA on Benchgen

No Benchgen results yet — be the first to run Harvey Lab-AA.

Harvey Lab-AA vs Other Benchmarks

BenchmarkWhat it testsSaturation
Harvey Lab-AALegal knowledge-work qualityLow
Legal Research BenchLegal research task accuracyLow
AA-BriefcaseProfessional knowledge-work qualityLow
CorpFin v2Corporate finance task automationLow

Run Harvey Lab-AA on Your Model

Benchgen lets you run Harvey Lab-AA against your own model, tracking legal knowledge-work accuracy over time as you evaluate deployment for legal-domain use cases.

Frequently Asked Questions

What is Harvey Lab-AA? Harvey Lab-AA is a legal knowledge-work benchmark testing AI models on legal research, analysis, and drafting quality.
What does a good score look like on Harvey Lab-AA? Kimi K3 reports 94.6% as of July 2026, a very strong result reflecting high accuracy on legal-domain tasks.
Who created Harvey Lab-AA? Harvey Lab-AA's originating team is not yet independently documented outside of its citation in Kimi K3's July 2026 technical report.