Benchgen

Multilingual MMBench — Results

RankModelScore
1north-micro-vision-instruct63.6
M

Multilingual MMBench

1 phaseActive

The multilingual development split of MMBench — the same 20 fine-grained VQA ability dimensions, evaluated across multiple languages.

Overview

Multilingual MMBench

Category Metric Saturation Created

Paper GitHub Dataset

Quick answer: Multilingual MMBench (MTL_MMBench_DEV) is the multilingual development split of the MMBench benchmark (Liu et al., 2023) — the same fine-grained VQA ability dimensions and CircularEval protocol as the English MMBench, evaluated across multiple non-English languages. North Micro Vision scores 63.6% as of August 2026.

At a Glance

What it tests: MMBench's ~20 fine-grained perception and reasoning ability dimensions, evaluated via questions and answer choices presented in multiple languages rather than English only.

Why it matters: Most vision-language benchmarks are English-centric; Multilingual MMBench surfaces the real capability gap many models show when the same visual reasoning task is posed in another language, which matters for any product serving a non-English-speaking user base.

Known limitations: Coverage and translation quality vary by language, and — like base MMBench — some ability dimensions are approaching saturation for larger frontier models.

What Multilingual MMBench Measures

Multilingual MMBench reuses MMBench's question bank and CircularEval scoring methodology (rotating multiple-choice answer order and requiring correctness across all rotations), but presents the questions and answer options in languages other than English. This isolates how much of a model's visual reasoning capability is coupled to English-language pretraining data versus genuinely transferable across languages.

For vision-language models trained with a deliberate multilingual data mixture — such as North Micro Vision, which reports a dedicated multilingual capability group in its evaluation — this split is a meaningful complement to MMMB, testing the same underlying fine-grained reasoning skills as MMBench rather than the broader open-ended multilingual VQA format MMMB uses.

Models with English-centric supervised fine-tuning data typically show a noticeable score drop moving from base MMBench to its multilingual split; smaller drops indicate more successful multilingual alignment during training.

Benchmark Specifications

FieldValue
Task categoryMultilingual (fine-grained VQA)
Metric% accuracy (CircularEval multiple-choice)
SaturationMedium
Created byLiu et al.
Source paperLiu et al. 2023
GitHubopen-compass/MMBench
Datasetlmms-lab/MMBench on HuggingFace

How Multilingual MMBench Is Scored

Scoring follows the same CircularEval protocol as base MMBench — a question counts as correct only if the model answers correctly across every cyclic shuffling of its multiple-choice options — applied to the multilingual development (MTL_MMBench_DEV) split.

State-of-the-Art Results

Score sourced from Cohere Labs' North Micro Vision Instruct launch announcement, August 2026, evaluated on the MTL_MMBench_DEV split.

Multilingual MMBench on Benchgen

No Benchgen results yet — be the first to run Multilingual MMBench.

Multilingual MMBench vs Other Benchmarks

BenchmarkWhat it testsSaturation
Multilingual MMBenchMMBench's fine-grained VQA dimensions, evaluated multilinguallyMedium
MMBenchThe same ability dimensions, English onlyMedium
MMMBBroader open-ended multilingual multimodal QA across 6 languagesLow

Use Multilingual MMBench alongside base MMBench to directly measure a model's English-to-multilingual capability gap on identical reasoning tasks; use MMMB for a broader, purpose-built multilingual benchmark.

Run Multilingual MMBench on Your Model

Benchgen lets you evaluate your own vision-language model's multilingual visual reasoning and compare it directly against its English-only MMBench score.

Frequently Asked Questions

What is Multilingual MMBench? Multilingual MMBench is the multilingual development split of MMBench, testing the same fine-grained VQA ability dimensions using non-English questions and answer choices.
What does a good Multilingual MMBench score look like? North Micro Vision Instruct reports 63.6% as of August 2026 — a modest drop from its 68.7% English MMBench score, indicating reasonable multilingual transfer.
Who created Multilingual MMBench? It uses the same benchmark methodology and question bank created by Liu et al. for MMBench; see the original paper.
Is Multilingual MMBench saturated? It shows medium saturation, similar to base MMBench, though it remains discriminative for models without dedicated multilingual training.