Benchgen

AA-LCR — Results

RankModelScore
1muse-glimmer80
2kimi-k374.7
3solar-pro-471
4nemotron-3-5-lightning-30b-a3b49.2
A

AA-LCR

1 phaseActive

Artificial Analysis's Long Context Reasoning benchmark, testing multi-step reasoning and synthesis across very long input documents. Metric: % accuracy.

Overview

AA-LCR

Category Metric Saturation Created

Website

Quick answer: AA-LCR (Artificial Analysis Long Context Reasoning) is Artificial Analysis's benchmark for evaluating multi-step reasoning and synthesis across very long input documents, part of their broader Intelligence Index suite used to independently compare frontier models. Kimi K3 scores 74.7% as of July 2026. Note: this is a distinct benchmark from Benchgen's general "Long Context Reasoning" page (based on Google DeepMind's Michelangelo benchmark).

At a Glance

What it tests: A model's ability to retain, retrieve, and reason across information scattered throughout very long input contexts, rather than relying on a single localized passage.

Why it matters: As context windows have grown into the hundreds of thousands or millions of tokens, raw retrieval is no longer the bottleneck — multi-hop reasoning across that context is. AA-LCR specifically targets this harder capability, and being run by an independent evaluator (Artificial Analysis) makes cross-model comparisons more trustworthy.

Known limitations: As with other Artificial Analysis benchmarks, detailed task construction methodology is published on their website rather than in an academic paper, so it should be treated as an industry-standard third-party evaluation rather than a peer-reviewed benchmark.

What AA-LCR Measures

AA-LCR evaluates a model's ability to perform multi-step reasoning and information synthesis across very long input documents — testing whether a model can connect facts, resolve references, and draw conclusions using information distributed across a large context window, rather than simply retrieving a single fact. It is maintained by Artificial Analysis as part of their independent, standardized model evaluation suite alongside benchmarks like their Intelligence Index and AA-Briefcase.

Benchmark Specifications

FieldValue
Task categoryReasoning / long context
Metric% accuracy
SaturationLow
Created byArtificial Analysis
Websiteartificialanalysis.ai

How AA-LCR Is Scored

Models answer questions or complete synthesis tasks that require reasoning across information spread throughout a long input document, scored on % accuracy against Artificial Analysis's reference answers.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K374.7%Kimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

AA-LCR on Benchgen

No Benchgen results yet — be the first to run AA-LCR.

AA-LCR vs Other Benchmarks

BenchmarkWhat it testsSaturation
AA-LCRLong-context multi-step reasoning (Artificial Analysis)Low
Long Context ReasoningLong-context reasoning (Michelangelo, Google DeepMind)Low
MRCR v2Multi-round co-reference resolutionLow
GDPval-AA v2Economically valuable task performance (Artificial Analysis)Low

Run AA-LCR on Your Model

Benchgen lets you run AA-LCR-style long-context reasoning evaluations against your own model, tracking accuracy as you scale context window usage in production.

Frequently Asked Questions

What is AA-LCR? AA-LCR (Artificial Analysis Long Context Reasoning) is a benchmark from Artificial Analysis testing multi-step reasoning and synthesis across very long input documents.
What does a good score look like on AA-LCR? Kimi K3 reports 74.7% as of July 2026, a strong result for long-context multi-hop reasoning among frontier models.
Who created AA-LCR? AA-LCR is created and maintained by Artificial Analysis, an independent AI model evaluation and analysis platform.
Is AA-LCR the same as Benchgen's "Long Context Reasoning" benchmark? No. Benchgen's separate "Long Context Reasoning" benchmark page is based on Google DeepMind's Michelangelo benchmark. AA-LCR is Artificial Analysis's own distinct long-context reasoning evaluation.