| Rank | Model | Score |
|---|---|---|
| 1 | claude-opus-5 | 1.18 |
| 2 | claude-fable-5 | 1.13 |
| 3 | grok-4-5 | 1.09 |
| 4 | claude-opus-4-8 | 1.05 |
| 5 | claude-sonnet-5 | 1.02 |
| 6 | gpt-5-6-sol | 0.9 |
| 7 | gpt-5-6-luna | 0.88 |
| 8 | gpt-5-6-terra | 0.85 |
1 phaseActive
Legora's Benchmark for Agentic Reasoning — 5,161+ real legal tasks across 28 practice areas, scored by weighted Quality Rubric Score inside the Legora harness.
Quick answer: The Legora Benchmark for Agentic Reasoning (BAR) evaluates AI models on 5,161+ real legal tasks across 28 practice areas — from M&A redlines to transfer pricing master files — using weighted Quality Rubric Scores inside the Legora agentic harness. Because cases are private and drawn from real law firm work, contamination is structurally prevented and scores reflect first-time performance, not pattern matching to training data.
What it tests: End-to-end legal work including document drafting, review, redlining, and advisory tasks, classified by difficulty (short/medium/long) across 28 practice areas, evaluated inside a full agentic harness with real matter file rooms.
Why it matters: Legal AI quality depends on the whole system — model, tools, harness, and retrieval — not the base model alone. Legora BAR is one of the few public benchmarks that measures the complete harness-model stack on genuine legal deliverables rather than synthetic simplified tasks.
Known limitations: Cases are private (not reproducible externally), all results are from the Legora harness only (scores are not portable to other harnesses), and absolute Quality Rubric Score percentages are not published — only relative performance indices comparing models evaluated in the same run.
Legora BAR tests whether an AI agent can complete the kinds of tasks that real legal teams assign every day. Each evaluation case, or eval, consists of four elements: a case description (the prompt a lawyer would write), a matter file room (the document folder with templates, precedents, and case materials), the system's output (one or more documents produced), and rubrics (binary, expert-written criteria graded by a legal professional). Practice area coverage spans M&A (850 cases), IP & tech (700), real estate (600), capital markets (275), litigation (286), and 23 additional areas. The benchmark uses 11,075+ source documents across 5,161 cases.
Difficulty is defined in terms of senior-associate effort: Short (a few hours to half a day — bounded, well-specified tasks), Medium (one to three days — a complete work product like a memo or redline), and Long (a week or more — full end-to-end matter delivery). Models perform similarly on short tasks and begin to diverge substantially on long tasks, where sustaining judgment across a large, messy document record is critical.
The private-case design is a deliberate architectural choice. Open-sourced benchmark cases routinely appear in the training data of the next generation of models, meaning high scores can reflect familiarity with the test rather than genuine legal capability. By keeping cases private and co-created with recognized law firm partners, Legora BAR ensures that every score represents a model encountering real work for the first time. One case — a fully synthetic transfer pricing Master File task co-created with AfterQuery — has been open-sourced for transparency and methodology review.
| Field | Value |
|---|---|
| Task category | Agent (legal) |
| Metric | Quality Rubric Index (relative, model average = 1.00) |
| Number of tasks | 5,161+ |
| Practice areas | 28 |
| Source documents | 11,075+ |
| Difficulty levels | Short · Medium · Long |
| Saturation | Low |
| Created by | Lauritzen, Sjölander, Helfer, Danil (Legora) |
| Technical report | Legora BAR – Introducing the Legora BAR |
| Open-source case | legora-oss/legora-bar-tax-case |
| Updated | July 28, 2026 |
Each evaluation case comes with a rubric: a list of binary, expert-written criteria covering facts, analysis, citations, and recommendations. The primary score is the Quality Rubric Score — the percentage of rubric criteria the agent satisfies. Criteria are weighted by importance (high / medium / low), so failing a high-importance criterion (e.g., missing a material change-of-control provision) penalises the score far more than failing a low-importance one (e.g., missing a secondary filing deadline). This weighted design means a materially incomplete deliverable still scores like one, rather than being masked by high coverage on minor points.
Legora runs each case three times in an isolated sandbox that mirrors the production client environment, then uses an LLM-as-a-judge to score each output against the rubric. The published leaderboard reports relative Quality Rubric Index values — how each model's score compares to the average of all models tested in the same evaluation run, where 1.00 equals the group average. Absolute percentage scores are not published externally. The benchmark also tracks two citation metrics: Cited Answers (whether each claim carries a citation) and Grounding (whether the sources the model consulted are reflected in its citations).
A 5% average increase in relative output quality was observed between June and July 2026 on the same set of models, attributable entirely to improvements in the Legora harness — demonstrating that the benchmark can detect harness-level regressions and improvements independently of model changes.
Scores are relative Quality Rubric Index values (average of all models in the evaluation = 1.00). Models run inside the Legora harness on the same private cases. Source: Legora BAR technical report, July 2026.
| Rank | Model | Quality Index (overall) | Best at |
|---|---|---|---|
| 1 | Claude Opus 5 | 1.18 | Long tasks, consistent across all difficulty levels |
| 2 | Claude Fable 5 | 1.13 | Long tasks, strongest grounding score |
| 3 | Grok 4.5 | 1.09 | Fastest latency at above-average quality |
| 4 | Claude Opus 4.8 | 1.05 | Balanced quality/cost |
| 5 | Claude Sonnet 5 | 1.02 | Near-average overall |
| 6 | GPT-5.6 Sol | 0.90 | Cited Answers score above average |
| 7 | GPT-5.6 Luna | 0.88 | Cited Answers score above average |
| 8 | GPT-5.6 Terra | 0.85 | Cited Answers score above average |
Scores are relative indices from Legora's published report. Absolute Quality Rubric Score percentages are not published externally. All models evaluated on the Legora harness — scores are not portable to other harnesses or prompt formats.
No Benchgen results yet — be the first to run Legora BAR.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| Legora BAR | End-to-end legal agent tasks (28 practice areas, real cases) | 5,161+ | Low |
| Harvey Lab AA | Legal agentic tasks via Harvey's harness | N/A | Low |
| Legal Research Bench | Legal research accuracy, citation quality | N/A | Low |
| Enterprise Claw Bench | Enterprise agent tasks (multi-domain) | 852 | Low |
Legora BAR is uniquely positioned among legal benchmarks: it is the only one to evaluate a full harness-model stack on real (private) law firm cases across 28 practice areas, with weighted rubric scoring that penalises high-severity misses more heavily than minor omissions. Teams that need to understand how an AI system performs on end-to-end legal deliverables — rather than discrete factual questions — should prioritise Legora BAR over narrower legal research benchmarks.
BenchGen lets teams track their agent's performance on benchmarks like Legora BAR across model versions and harness changes — turning a one-time vendor report into a continuous regression signal. Instead of relying on Legora's published snapshot, you can connect your own agent, run the open-sourced transfer pricing case as a reproducible eval, and compare results across every configuration change. Start tracking on BenchGen.
Benchmark definition paraphrased from Lauritzen et al. 2026. State-of-the-art scores sourced from Legora's published technical report (July 28, 2026) and reported as relative Quality Rubric Index values. Last updated 2026-08-03.