Benchgen

MMMB — Results

RankModelScore
1north-micro-vision-instruct72.8
M

MMMB

1 phaseActive

6-language, 15-category, 12,000-question multimodal benchmark from the Parrot paper, testing genuine multilingual visual understanding.

Overview

MMMB

Category Metric Tasks Saturation Created

Paper GitHub

Quick answer: MMMB (the Massive Multilingual Multimodal Benchmark, introduced in Sun et al.'s "Parrot" paper, 2024) evaluates vision-language models across 6 languages, 15 categories, and 12,000 questions, purpose-built to measure genuine multilingual multimodal understanding rather than English performance with translated prompts. North Micro Vision scores 72.8% as of August 2026.

At a Glance

What it tests: Multimodal question answering across 6 languages and 15 question categories (12,000 questions total), designed to reveal whether a model's visual reasoning transfers across languages or degrades outside English.

Why it matters: The Parrot paper showed that supervised fine-tuning on English-centric multimodal instruction data often erodes a model's multilingual capability as training progresses; MMMB was built specifically to quantify and track this degradation.

Known limitations: Being purpose-built and relatively new, MMMB has less independent third-party adoption than older benchmarks like MMBench, though it remains one of the more rigorous multilingual multimodal benchmarks available.

What MMMB Measures

MMMB was introduced alongside PARROT, a method for aligning multilingual visual tokens using textual guidance and mixture-of-experts routing. The benchmark's authors observed that imbalanced, largely English-centric supervised fine-tuning data causes many multimodal LLMs to lose multilingual capability as training progresses — a failure mode standard English-only benchmarks can't detect.

MMMB addresses this by spanning 6 languages, 15 question categories, and 12,000 total questions, giving broad category and language coverage rather than a narrow translated subset. This makes it a purpose-built complement to translated splits of English-first benchmarks (like Multilingual MMBench), since it was designed from the ground up as a multilingual benchmark rather than adapted from an English original.

For models like North Micro Vision that report a dedicated "Multilingual" capability group, MMMB is typically the stronger of the two multilingual benchmarks (alongside Multilingual MMBench), since its broader category and language coverage makes it harder to do well on through narrow multilingual fine-tuning alone.

Benchmark Specifications

FieldValue
Task categoryMultilingual (multimodal QA)
Metric% accuracy
Number of tasks12,000 (6 languages x 15 categories)
SaturationLow
Created bySun et al.
Source paperSun et al. 2024
GitHubAIDC-AI/Parrot

How MMMB Is Scored

Models answer multiple-choice or short-answer multimodal questions across all 6 languages and 15 categories; the overall score is percent accuracy averaged across the full question set, with per-language breakdowns also commonly reported.

State-of-the-Art Results

Score sourced from Cohere Labs' North Micro Vision Instruct launch announcement, August 2026.

MMMB on Benchgen

No Benchgen results yet — be the first to run MMMB.

MMMB vs Other Benchmarks

BenchmarkWhat it testsSaturation
MMMBPurpose-built multilingual multimodal QA, 6 languagesLow
Multilingual MMBenchMMBench's fine-grained VQA dimensions, translatedMedium
MMBenchFine-grained English-only VQA ability dimensionsMedium

MMMB and Multilingual MMBench are complementary: MMMB was designed as a multilingual benchmark from the ground up, while Multilingual MMBench directly measures the drop from an identical English-only baseline.

Run MMMB on Your Model

Benchgen lets you evaluate your own vision-language model's multilingual multimodal understanding across 6 languages and track results across model versions.

Frequently Asked Questions

What is MMMB? MMMB (Massive Multilingual Multimodal Benchmark) is a purpose-built multilingual benchmark spanning 6 languages, 15 categories, and 12,000 questions, introduced in the Parrot paper.
What does a good MMMB score look like? North Micro Vision Instruct reports 72.8% as of August 2026; scores vary significantly by language and model multilingual training coverage.
Who created MMMB? MMMB was introduced by Sun et al. in the PARROT paper; see the original paper.
Is MMMB saturated? No — as a newer, purpose-built multilingual benchmark, MMMB remains low-saturation and discriminative across model sizes.