Benchgen

Legora BAR — Results

RankModelScore
1claude-opus-51.18
2claude-fable-51.13
3grok-4-51.09
4claude-opus-4-81.05
5claude-sonnet-51.02
6gpt-5-6-sol0.9
7gpt-5-6-luna0.88
8gpt-5-6-terra0.85
L

Legora BAR

1 phaseActive

Legora's Benchmark for Agentic Reasoning — 5,161+ real legal tasks across 28 practice areas, scored by weighted Quality Rubric Score inside the Legora harness.

Overview

Legora BAR

Category Metric Tasks Saturation Created

Technical Report GitHub

Quick answer: The Legora Benchmark for Agentic Reasoning (BAR) evaluates AI models on 5,161+ real legal tasks across 28 practice areas — from M&A redlines to transfer pricing master files — using weighted Quality Rubric Scores inside the Legora agentic harness. Because cases are private and drawn from real law firm work, contamination is structurally prevented and scores reflect first-time performance, not pattern matching to training data.

At a Glance

What it tests: End-to-end legal work including document drafting, review, redlining, and advisory tasks, classified by difficulty (short/medium/long) across 28 practice areas, evaluated inside a full agentic harness with real matter file rooms.

Why it matters: Legal AI quality depends on the whole system — model, tools, harness, and retrieval — not the base model alone. Legora BAR is one of the few public benchmarks that measures the complete harness-model stack on genuine legal deliverables rather than synthetic simplified tasks.

Known limitations: Cases are private (not reproducible externally), all results are from the Legora harness only (scores are not portable to other harnesses), and absolute Quality Rubric Score percentages are not published — only relative performance indices comparing models evaluated in the same run.

What Legora BAR Measures

Legora BAR tests whether an AI agent can complete the kinds of tasks that real legal teams assign every day. Each evaluation case, or eval, consists of four elements: a case description (the prompt a lawyer would write), a matter file room (the document folder with templates, precedents, and case materials), the system's output (one or more documents produced), and rubrics (binary, expert-written criteria graded by a legal professional). Practice area coverage spans M&A (850 cases), IP & tech (700), real estate (600), capital markets (275), litigation (286), and 23 additional areas. The benchmark uses 11,075+ source documents across 5,161 cases.

Difficulty is defined in terms of senior-associate effort: Short (a few hours to half a day — bounded, well-specified tasks), Medium (one to three days — a complete work product like a memo or redline), and Long (a week or more — full end-to-end matter delivery). Models perform similarly on short tasks and begin to diverge substantially on long tasks, where sustaining judgment across a large, messy document record is critical.

The private-case design is a deliberate architectural choice. Open-sourced benchmark cases routinely appear in the training data of the next generation of models, meaning high scores can reflect familiarity with the test rather than genuine legal capability. By keeping cases private and co-created with recognized law firm partners, Legora BAR ensures that every score represents a model encountering real work for the first time. One case — a fully synthetic transfer pricing Master File task co-created with AfterQuery — has been open-sourced for transparency and methodology review.

Benchmark Specifications

FieldValue
Task categoryAgent (legal)
MetricQuality Rubric Index (relative, model average = 1.00)
Number of tasks5,161+
Practice areas28
Source documents11,075+
Difficulty levelsShort · Medium · Long
SaturationLow
Created byLauritzen, Sjölander, Helfer, Danil (Legora)
Technical reportLegora BAR – Introducing the Legora BAR
Open-source caselegora-oss/legora-bar-tax-case
UpdatedJuly 28, 2026

How Legora BAR Is Scored

Each evaluation case comes with a rubric: a list of binary, expert-written criteria covering facts, analysis, citations, and recommendations. The primary score is the Quality Rubric Score — the percentage of rubric criteria the agent satisfies. Criteria are weighted by importance (high / medium / low), so failing a high-importance criterion (e.g., missing a material change-of-control provision) penalises the score far more than failing a low-importance one (e.g., missing a secondary filing deadline). This weighted design means a materially incomplete deliverable still scores like one, rather than being masked by high coverage on minor points.

Legora runs each case three times in an isolated sandbox that mirrors the production client environment, then uses an LLM-as-a-judge to score each output against the rubric. The published leaderboard reports relative Quality Rubric Index values — how each model's score compares to the average of all models tested in the same evaluation run, where 1.00 equals the group average. Absolute percentage scores are not published externally. The benchmark also tracks two citation metrics: Cited Answers (whether each claim carries a citation) and Grounding (whether the sources the model consulted are reflected in its citations).

A 5% average increase in relative output quality was observed between June and July 2026 on the same set of models, attributable entirely to improvements in the Legora harness — demonstrating that the benchmark can detect harness-level regressions and improvements independently of model changes.

State-of-the-Art Results

Scores are relative Quality Rubric Index values (average of all models in the evaluation = 1.00). Models run inside the Legora harness on the same private cases. Source: Legora BAR technical report, July 2026.

RankModelQuality Index (overall)Best at
1Claude Opus 51.18Long tasks, consistent across all difficulty levels
2Claude Fable 51.13Long tasks, strongest grounding score
3Grok 4.51.09Fastest latency at above-average quality
4Claude Opus 4.81.05Balanced quality/cost
5Claude Sonnet 51.02Near-average overall
6GPT-5.6 Sol0.90Cited Answers score above average
7GPT-5.6 Luna0.88Cited Answers score above average
8GPT-5.6 Terra0.85Cited Answers score above average

Scores are relative indices from Legora's published report. Absolute Quality Rubric Score percentages are not published externally. All models evaluated on the Legora harness — scores are not portable to other harnesses or prompt formats.

Legora BAR on Benchgen

No Benchgen results yet — be the first to run Legora BAR.

Legora BAR vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
Legora BAREnd-to-end legal agent tasks (28 practice areas, real cases)5,161+Low
Harvey Lab AALegal agentic tasks via Harvey's harnessN/ALow
Legal Research BenchLegal research accuracy, citation qualityN/ALow
Enterprise Claw BenchEnterprise agent tasks (multi-domain)852Low

Legora BAR is uniquely positioned among legal benchmarks: it is the only one to evaluate a full harness-model stack on real (private) law firm cases across 28 practice areas, with weighted rubric scoring that penalises high-severity misses more heavily than minor omissions. Teams that need to understand how an AI system performs on end-to-end legal deliverables — rather than discrete factual questions — should prioritise Legora BAR over narrower legal research benchmarks.

Run Legora BAR on Your Model

BenchGen lets teams track their agent's performance on benchmarks like Legora BAR across model versions and harness changes — turning a one-time vendor report into a continuous regression signal. Instead of relying on Legora's published snapshot, you can connect your own agent, run the open-sourced transfer pricing case as a reproducible eval, and compare results across every configuration change. Start tracking on BenchGen.

Frequently Asked Questions

What is Legora BAR? Legora BAR (Benchmark for Agentic Reasoning) is a private benchmark created by Legora that evaluates AI systems on over 5,000 real legal tasks across 28 practice areas, scoring outputs using expert-written weighted rubrics inside the Legora agentic harness.
What does a good Legora BAR score look like? Scores are reported as a relative Quality Rubric Index where 1.00 equals the average of all models tested. As of July 2026, the top performer (Claude Opus 5) achieved a relative index of approximately 1.18, while the lowest-ranked model (GPT-5.6 Terra) scored approximately 0.85. Models above 1.05 are performing meaningfully better than average on real legal work.
Who created Legora BAR? Legora BAR was created by Jacob Lauritzen, Emil Sjölander, Ebba Helfer, and Firas Danil at Legora, in collaboration with law firm partners and AfterQuery. The technical report was published July 28, 2026. See: legora.com/bar.
Is Legora BAR saturated? Saturation is low. Cases are drawn from real law firm work and kept private, structurally preventing training-data contamination. The hardest long-form cases — full end-to-end matter deliveries — are still being added to the corpus as models become capable enough for the measurement to be meaningful.
How does Legora BAR differ from other legal AI benchmarks? Most legal benchmarks test discrete knowledge retrieval or simplified tasks in synthetic environments. Legora BAR evaluates the complete harness-model stack on end-to-end legal deliverables in a production-equivalent environment, with private cases to prevent contamination and weighted rubric scoring that mirrors how senior lawyers actually review work.

Benchmark definition paraphrased from Lauritzen et al. 2026. State-of-the-art scores sourced from Legora's published technical report (July 28, 2026) and reported as relative Quality Rubric Index values. Last updated 2026-08-03.