| Rank | Model | Score |
|---|---|---|
| 1 | north-micro-vision-instruct | 72.8 |
1 phaseActive
6-language, 15-category, 12,000-question multimodal benchmark from the Parrot paper, testing genuine multilingual visual understanding.
Quick answer: MMMB (the Massive Multilingual Multimodal Benchmark, introduced in Sun et al.'s "Parrot" paper, 2024) evaluates vision-language models across 6 languages, 15 categories, and 12,000 questions, purpose-built to measure genuine multilingual multimodal understanding rather than English performance with translated prompts. North Micro Vision scores 72.8% as of August 2026.
What it tests: Multimodal question answering across 6 languages and 15 question categories (12,000 questions total), designed to reveal whether a model's visual reasoning transfers across languages or degrades outside English.
Why it matters: The Parrot paper showed that supervised fine-tuning on English-centric multimodal instruction data often erodes a model's multilingual capability as training progresses; MMMB was built specifically to quantify and track this degradation.
Known limitations: Being purpose-built and relatively new, MMMB has less independent third-party adoption than older benchmarks like MMBench, though it remains one of the more rigorous multilingual multimodal benchmarks available.
MMMB was introduced alongside PARROT, a method for aligning multilingual visual tokens using textual guidance and mixture-of-experts routing. The benchmark's authors observed that imbalanced, largely English-centric supervised fine-tuning data causes many multimodal LLMs to lose multilingual capability as training progresses — a failure mode standard English-only benchmarks can't detect.
MMMB addresses this by spanning 6 languages, 15 question categories, and 12,000 total questions, giving broad category and language coverage rather than a narrow translated subset. This makes it a purpose-built complement to translated splits of English-first benchmarks (like Multilingual MMBench), since it was designed from the ground up as a multilingual benchmark rather than adapted from an English original.
For models like North Micro Vision that report a dedicated "Multilingual" capability group, MMMB is typically the stronger of the two multilingual benchmarks (alongside Multilingual MMBench), since its broader category and language coverage makes it harder to do well on through narrow multilingual fine-tuning alone.
| Field | Value |
|---|---|
| Task category | Multilingual (multimodal QA) |
| Metric | % accuracy |
| Number of tasks | 12,000 (6 languages x 15 categories) |
| Saturation | Low |
| Created by | Sun et al. |
| Source paper | Sun et al. 2024 |
| GitHub | AIDC-AI/Parrot |
Models answer multiple-choice or short-answer multimodal questions across all 6 languages and 15 categories; the overall score is percent accuracy averaged across the full question set, with per-language breakdowns also commonly reported.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | North Micro Vision Instruct | 72.8% | North Micro Vision launch blog | 2026-08 |
Score sourced from Cohere Labs' North Micro Vision Instruct launch announcement, August 2026.
No Benchgen results yet — be the first to run MMMB.
| Benchmark | What it tests | Saturation |
|---|---|---|
| MMMB | Purpose-built multilingual multimodal QA, 6 languages | Low |
| Multilingual MMBench | MMBench's fine-grained VQA dimensions, translated | Medium |
| MMBench | Fine-grained English-only VQA ability dimensions | Medium |
MMMB and Multilingual MMBench are complementary: MMMB was designed as a multilingual benchmark from the ground up, while Multilingual MMBench directly measures the drop from an identical English-only baseline.
Benchgen lets you evaluate your own vision-language model's multilingual multimodal understanding across 6 languages and track results across model versions.