| Rank | Model | Score |
|---|---|---|
| 1 | muse-glimmer | 80 |
| 2 | kimi-k3 | 74.7 |
| 3 | solar-pro-4 | 71 |
| 4 | nemotron-3-5-lightning-30b-a3b | 49.2 |
1 phaseActive
Artificial Analysis's Long Context Reasoning benchmark, testing multi-step reasoning and synthesis across very long input documents. Metric: % accuracy.
Quick answer: AA-LCR (Artificial Analysis Long Context Reasoning) is Artificial Analysis's benchmark for evaluating multi-step reasoning and synthesis across very long input documents, part of their broader Intelligence Index suite used to independently compare frontier models. Kimi K3 scores 74.7% as of July 2026. Note: this is a distinct benchmark from Benchgen's general "Long Context Reasoning" page (based on Google DeepMind's Michelangelo benchmark).
What it tests: A model's ability to retain, retrieve, and reason across information scattered throughout very long input contexts, rather than relying on a single localized passage.
Why it matters: As context windows have grown into the hundreds of thousands or millions of tokens, raw retrieval is no longer the bottleneck — multi-hop reasoning across that context is. AA-LCR specifically targets this harder capability, and being run by an independent evaluator (Artificial Analysis) makes cross-model comparisons more trustworthy.
Known limitations: As with other Artificial Analysis benchmarks, detailed task construction methodology is published on their website rather than in an academic paper, so it should be treated as an industry-standard third-party evaluation rather than a peer-reviewed benchmark.
AA-LCR evaluates a model's ability to perform multi-step reasoning and information synthesis across very long input documents — testing whether a model can connect facts, resolve references, and draw conclusions using information distributed across a large context window, rather than simply retrieving a single fact. It is maintained by Artificial Analysis as part of their independent, standardized model evaluation suite alongside benchmarks like their Intelligence Index and AA-Briefcase.
| Field | Value |
|---|---|
| Task category | Reasoning / long context |
| Metric | % accuracy |
| Saturation | Low |
| Created by | Artificial Analysis |
| Website | artificialanalysis.ai |
Models answer questions or complete synthesis tasks that require reasoning across information spread throughout a long input document, scored on % accuracy against Artificial Analysis's reference answers.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 74.7% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run AA-LCR.
| Benchmark | What it tests | Saturation |
|---|---|---|
| AA-LCR | Long-context multi-step reasoning (Artificial Analysis) | Low |
| Long Context Reasoning | Long-context reasoning (Michelangelo, Google DeepMind) | Low |
| MRCR v2 | Multi-round co-reference resolution | Low |
| GDPval-AA v2 | Economically valuable task performance (Artificial Analysis) | Low |
Benchgen lets you run AA-LCR-style long-context reasoning evaluations against your own model, tracking accuracy as you scale context window usage in production.