Benchgen

MultiChallenge — Results

RankModelScore
1nova-2-pro0.777
2nova-2-lite0.766
3nova-2-omni0.755
4gpt-50.696
5qwen3-5-397b-a17b0.676
6nemotron-3-ultra-550b-a55b0.638
7qwen3-5-122b-a10b0.615
8qwen3-5-27b0.608
9o30.604
10qwen3-5-35b-a3b0.6
11nemotron-3-super-120b-a12b0.552
12qwen3-5-9b0.545
13kimi-k2-instruct-09050.541
14kimi-k2-instruct0.541
15mai-thinking-10.53
16qwen3-5-4b0.49
17minimax-m1-40k0.447
18minimax-m1-80k0.447
19gpt-4-50.438
20o4-mini0.43
21gpt-4o0.403
22o3-mini0.399
23nemotron-3-nano-30b-a3b0.385
24gpt-4-10.383
25gpt-4-1-mini0.358

MultiChallenge

1 phaseActive

Realistic multi-turn conversation benchmark testing instruction retention, inference memory, versioned editing, and self-coherence across 29 frontier models. Metric: accuracy.

Overview

MultiChallenge

Category Metric Saturation Modality

Paper GitHub

Quick answer: MultiChallenge is a realistic multi-turn conversation benchmark by Sirdeshmukh et al. (2025) that evaluates LLMs on four categories: instruction retention, inference memory, reliable versioned editing, and self-coherence. Nova 2 Pro leads with 77.7% across 29 evaluated models as of August 2026.


What Does MultiChallenge Test?

MultiChallenge evaluates models on sustained, contextually complex dialogues — the kind of conversations that arise in travel planning, technical documentation, and professional communication. Unlike single-turn benchmarks, it requires models to maintain consistency and follow evolving instructions across multiple conversation turns.

CategoryWhat it tests
Instruction retentionMaintain instructions given at the start throughout a long conversation
Inference memoryRecall and connect details from earlier turns to answer later questions
Reliable versioned editingAdapt correctly when instructions evolve during collaborative editing
Self-coherenceAvoid contradicting prior responses within the same conversation

How Is MultiChallenge Scored?

Each conversation is evaluated for correctness across all four categories. Scores are normalized to 0–1. The benchmark focuses on failure modes that appear only in multi-turn settings and are invisible to single-turn evaluations.


Key Facts

PropertyValue
PublishedJanuary 2025
MetricAccuracy
Score range0–1
Top modelNova 2 Pro (0.777)
Models evaluated29

FAQ

What is MultiChallenge? MultiChallenge is a realistic multi-turn conversation evaluation benchmark that tests LLMs on instruction retention, inference memory, versioned editing, and self-coherence across diverse dialogue scenarios.

Who created MultiChallenge? MultiChallenge was created by Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols, Lifeng Jin and colleagues at Amazon, published in January 2025 (arXiv 2501.17399).

How is MultiChallenge different from single-turn benchmarks? Most benchmarks evaluate isolated prompts. MultiChallenge tests whether models can maintain consistency and follow instructions across an entire multi-turn conversation — a key requirement for real-world assistant applications.

What score does the best model achieve on MultiChallenge? Nova 2 Pro from Amazon achieves 0.777 (77.7%), followed by Nova 2 Lite at 0.766 and Nova 2 Omni at 0.755.