| Rank | Model | Score |
|---|---|---|
| 1 | north-micro-vision-instruct | 63.6 |
1 phaseActive
The multilingual development split of MMBench — the same 20 fine-grained VQA ability dimensions, evaluated across multiple languages.
Quick answer: Multilingual MMBench (
MTL_MMBench_DEV) is the multilingual development split of the MMBench benchmark (Liu et al., 2023) — the same fine-grained VQA ability dimensions and CircularEval protocol as the English MMBench, evaluated across multiple non-English languages. North Micro Vision scores 63.6% as of August 2026.
What it tests: MMBench's ~20 fine-grained perception and reasoning ability dimensions, evaluated via questions and answer choices presented in multiple languages rather than English only.
Why it matters: Most vision-language benchmarks are English-centric; Multilingual MMBench surfaces the real capability gap many models show when the same visual reasoning task is posed in another language, which matters for any product serving a non-English-speaking user base.
Known limitations: Coverage and translation quality vary by language, and — like base MMBench — some ability dimensions are approaching saturation for larger frontier models.
Multilingual MMBench reuses MMBench's question bank and CircularEval scoring methodology (rotating multiple-choice answer order and requiring correctness across all rotations), but presents the questions and answer options in languages other than English. This isolates how much of a model's visual reasoning capability is coupled to English-language pretraining data versus genuinely transferable across languages.
For vision-language models trained with a deliberate multilingual data mixture — such as North Micro Vision, which reports a dedicated multilingual capability group in its evaluation — this split is a meaningful complement to MMMB, testing the same underlying fine-grained reasoning skills as MMBench rather than the broader open-ended multilingual VQA format MMMB uses.
Models with English-centric supervised fine-tuning data typically show a noticeable score drop moving from base MMBench to its multilingual split; smaller drops indicate more successful multilingual alignment during training.
| Field | Value |
|---|---|
| Task category | Multilingual (fine-grained VQA) |
| Metric | % accuracy (CircularEval multiple-choice) |
| Saturation | Medium |
| Created by | Liu et al. |
| Source paper | Liu et al. 2023 |
| GitHub | open-compass/MMBench |
| Dataset | lmms-lab/MMBench on HuggingFace |
Scoring follows the same CircularEval protocol as base MMBench — a question counts as correct only if the model answers correctly across every cyclic shuffling of its multiple-choice options — applied to the multilingual development (MTL_MMBench_DEV) split.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | North Micro Vision Instruct | 63.6% | North Micro Vision launch blog | 2026-08 |
Score sourced from Cohere Labs' North Micro Vision Instruct launch announcement, August 2026, evaluated on the MTL_MMBench_DEV split.
No Benchgen results yet — be the first to run Multilingual MMBench.
| Benchmark | What it tests | Saturation |
|---|---|---|
| Multilingual MMBench | MMBench's fine-grained VQA dimensions, evaluated multilingually | Medium |
| MMBench | The same ability dimensions, English only | Medium |
| MMMB | Broader open-ended multilingual multimodal QA across 6 languages | Low |
Use Multilingual MMBench alongside base MMBench to directly measure a model's English-to-multilingual capability gap on identical reasoning tasks; use MMMB for a broader, purpose-built multilingual benchmark.
Benchgen lets you evaluate your own vision-language model's multilingual visual reasoning and compare it directly against its English-only MMBench score.