Benchgen

Global-MMLU-Lite — Results

RankModelScore
1claude-fable-593.3
2gemini-3-1-pro92.7
3gpt-5-6-sol91.8
4deepseek-v4-pro89.3
5glm-5-289.2
6inkling88.7
7gemini-2-5-pro88.6
8kimi-k2-688.4
9nemotron-3-ultra-550b-a55b85.6
10kimi-k2-584
11gemini-2-0-flash-lite78.2
12gemma-3-27b75.1
13gemma-3-12b69.5
14gemma-3n-e4b-litert-preview64.5
15gemma-3n-e4b64.5
16gemma-3n-e2b-litert-preview59
17gemma-3n-e2b59
18gemma-3-4b54.5
19gemma-3-1b34.2

Global-MMLU-Lite

1 phaseActive

Multilingual MMLU variant from Cohere Labs — 9,200 human-translated 4-choice questions across 23 languages with cultural sensitivity annotations. Metric: % accuracy.

Overview

Global-MMLU-Lite

Category Metric Tasks Languages Saturation Created

Paper Dataset

Quick answer: Global-MMLU-Lite is a multilingual evaluation benchmark by Cohere Labs (Singh et al., 2024) that tests AI models on broad academic knowledge across 23 languages using 9,200 human-translated 4-choice questions drawn from the original MMLU dataset. It annotates questions as Culturally Sensitive or Culturally Agnostic to surface linguistic and cultural biases. Claude Fable 5 leads at 93.3% with GPT-5.6 Sol close behind at 91.8%.

At a Glance

What it tests: Multilingual academic knowledge across 23 languages — the same MMLU subject domains (STEM, humanities, social sciences, etc.) evaluated in the language the model is expected to serve, not just in English.

Why it matters: Most benchmark evaluations are English-only, masking large capability gaps in multilingual models. Global-MMLU-Lite makes language coverage measurable and comparable: a model that scores 92% in English but 65% in Arabic reveals a practical deployment problem that English-only benchmarks would miss.

Known limitations: Scores are averaged across languages; per-language breakdown may reveal significant variance. The 4-choice format inherited from MMLU carries some saturation risk for frontier models. Cultural annotations are based on regional collaborator judgments and may not be exhaustive.

What Global-MMLU-Lite Measures

Global-MMLU-Lite takes the original MMLU question bank — 14,000 questions across academic disciplines — and provides human-translated, professionally verified versions in 23 languages. For each language, the dataset includes 200 Culturally Sensitive (CS) questions (where the correct answer or framing may vary by cultural context) and 200 Culturally Agnostic (CA) questions (where the correct answer is universal). This 400-question-per-language structure makes cross-language comparisons fair and balanced.

The "Lite" designation distinguishes it from the larger Global-MMLU dataset (42 languages, 590K rows), which uses a mix of machine translations alongside professional ones. Global-MMLU-Lite is fully human-translated or post-edited — making it higher fidelity for evaluation, though more limited in language count.

Scores are reported as overall accuracy (% correct) averaged across all languages and both cultural subsets. Labs may also report per-language or per-subset breakdowns to diagnose specific gaps.

Benchmark Specifications

FieldValue
Task categoryReasoning / multilingual knowledge
Metric% accuracy (4-choice MCQ)
Number of tasks9,200 test samples (400 per language × 23 languages)
Languages23 (v3.0, May 2026)
Cultural subsetsCulturally Sensitive (CS) · Culturally Agnostic (CA)
SaturationMedium
Created byShivalika Singh, Angelika Romanou, Clémentine Fourrier, et al.
AffiliationCohere Labs
LicenseApache 2.0
First releaseDecember 2024 (v1.0, 15 languages)
Current version3.0 (May 2026, 23 languages)
Source paperGlobal MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation (arXiv 2412.03304)
DatasetCohereLabs/Global-MMLU-Lite on HuggingFace

Languages Covered (v3.0)

Arabic, Bengali, Chinese (Simplified), Czech, Dutch, French, German, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Oriya, Polish, Portuguese, Russian, Slovak, Spanish, Swahili, Tajik, Ukrainian, Welsh (+ English for CA/CS annotation reference).

How Global-MMLU-Lite Is Scored

A model is presented with a question in a target language and must select the correct answer from four options (A–D). Scores are reported as accuracy (% correct). The benchmark is typically evaluated in a 0-shot or 5-shot setting. Reported scores often reflect the average across all 23 languages; per-language breakdowns are available from the dataset but less commonly reported in model cards.

Frontier models tend to score in the high-80s to low-90s on Global-MMLU-Lite as of mid-2026, with closed-weight models leading. Performance gaps between English and low-resource languages (e.g. Swahili, Oriya) are typically 10–20 percentage points even for frontier models.

State-of-the-Art Results

Scores from Inkling model card (Thinking Machines Lab, July 2026), evaluated at effort=0.99.

RankModelScoreWeights
1Claude Fable 593.3%Closed
2GPT-5.6 Sol91.8%Closed
3Gemini 3.1 Pro92.7%Closed
4GLM 5.289.2%Open
5DeepSeek V4 Pro89.3%Open
6Inkling88.7%Open
7Kimi K2.688.4%Open
8Nemotron 3 Ultra85.6%Open
9Kimi K2.584.0%Open
BenchmarkLanguagesTasksCultural annotationsSaturation
Global-MMLU-Lite239,200Yes (CS / CA)Medium
MMLU-ProEnglish only12,032NoMedium
Global-MMLU (full)42~590KYesMedium

Last updated 2026-07-16.