| Rank | Model | Score |
|---|---|---|
| 1 | nova-2-pro | 0.777 |
| 2 | nova-2-lite | 0.766 |
| 3 | nova-2-omni | 0.755 |
| 4 | gpt-5 | 0.696 |
| 5 | qwen3-5-397b-a17b | 0.676 |
| 6 | nemotron-3-ultra-550b-a55b | 0.638 |
| 7 | qwen3-5-122b-a10b | 0.615 |
| 8 | qwen3-5-27b | 0.608 |
| 9 | o3 | 0.604 |
| 10 | qwen3-5-35b-a3b | 0.6 |
| 11 | nemotron-3-super-120b-a12b | 0.552 |
| 12 | qwen3-5-9b | 0.545 |
| 13 | kimi-k2-instruct-0905 | 0.541 |
| 14 | kimi-k2-instruct | 0.541 |
| 15 | mai-thinking-1 | 0.53 |
| 16 | qwen3-5-4b | 0.49 |
| 17 | minimax-m1-40k | 0.447 |
| 18 | minimax-m1-80k | 0.447 |
| 19 | gpt-4-5 | 0.438 |
| 20 | o4-mini | 0.43 |
| 21 | gpt-4o | 0.403 |
| 22 | o3-mini | 0.399 |
| 23 | nemotron-3-nano-30b-a3b | 0.385 |
| 24 | gpt-4-1 | 0.383 |
| 25 | gpt-4-1-mini | 0.358 |
1 phaseActive
Realistic multi-turn conversation benchmark testing instruction retention, inference memory, versioned editing, and self-coherence across 29 frontier models. Metric: accuracy.
Quick answer: MultiChallenge is a realistic multi-turn conversation benchmark by Sirdeshmukh et al. (2025) that evaluates LLMs on four categories: instruction retention, inference memory, reliable versioned editing, and self-coherence. Nova 2 Pro leads with 77.7% across 29 evaluated models as of August 2026.
MultiChallenge evaluates models on sustained, contextually complex dialogues — the kind of conversations that arise in travel planning, technical documentation, and professional communication. Unlike single-turn benchmarks, it requires models to maintain consistency and follow evolving instructions across multiple conversation turns.
| Category | What it tests |
|---|---|
| Instruction retention | Maintain instructions given at the start throughout a long conversation |
| Inference memory | Recall and connect details from earlier turns to answer later questions |
| Reliable versioned editing | Adapt correctly when instructions evolve during collaborative editing |
| Self-coherence | Avoid contradicting prior responses within the same conversation |
Each conversation is evaluated for correctness across all four categories. Scores are normalized to 0–1. The benchmark focuses on failure modes that appear only in multi-turn settings and are invisible to single-turn evaluations.
| Property | Value |
|---|---|
| Published | January 2025 |
| Metric | Accuracy |
| Score range | 0–1 |
| Top model | Nova 2 Pro (0.777) |
| Models evaluated | 29 |
What is MultiChallenge? MultiChallenge is a realistic multi-turn conversation evaluation benchmark that tests LLMs on instruction retention, inference memory, versioned editing, and self-coherence across diverse dialogue scenarios.
Who created MultiChallenge? MultiChallenge was created by Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols, Lifeng Jin and colleagues at Amazon, published in January 2025 (arXiv 2501.17399).
How is MultiChallenge different from single-turn benchmarks? Most benchmarks evaluate isolated prompts. MultiChallenge tests whether models can maintain consistency and follow instructions across an entire multi-turn conversation — a key requirement for real-world assistant applications.
What score does the best model achieve on MultiChallenge? Nova 2 Pro from Amazon achieves 0.777 (77.7%), followed by Nova 2 Lite at 0.766 and Nova 2 Omni at 0.755.