Benchgen

WMT25 — Results

RankModelScore
1hy-mt2-7b63.86
2hy-mt2-30b-a3b62.89
3hy-mt2-1-8b50.3
W

WMT25

1 phaseActive

WMT's annual General Machine Translation shared task — human-evaluation test sets, 12 translation directions. Metric: XCOMET-XXL (0-100).

Overview

WMT25

Category Metric Directions Saturation

Paper Shared Task

Quick answer: WMT25 is the 2025 edition of the Conference on Machine Translation's (WMT) annual General Machine Translation shared task — a harder, human-curated test set explicitly designed to resist saturation ("Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets"), covering 12 translation directions. It's the academic community's flagship head-to-head MT evaluation, run fresh each year against new source text.

At a Glance

What it tests: Translation quality on freshly-curated, deliberately-difficult source text across 12 language directions, evaluated by both automatic MT metrics and human annotators. Why it matters: Unlike static benchmarks that saturate over time, WMT refreshes its test set every year specifically to stay hard — the 2025 edition's own title flags this as an explicit design goal. Known limitations: Only 12 directions per edition (versus FLORES-200's 1,056), and coverage skews toward higher-resource language pairs relevant to the shared task's participant pool.

What WMT25 Measures

The Conference on Machine Translation (WMT) has run an annual General Machine Translation shared task since 2006, each year assembling a fresh test set of untranslated source documents and collecting both automatic-metric and human-evaluation judgments of submitted systems' translations. The WMT25 edition explicitly responds to benchmark saturation in prior years by curating harder source material, aiming to keep the shared task discriminative between top-tier systems.

WMT25 evaluates across 12 translation directions per the shared task's own scope. As with FLORES-200, Benchgen tracks the reference-based XCOMET-XXL metric, since it's the metric most consistently reported by LLM technical reports citing WMT25 results (papers frequently additionally report CometKiwi and GEMBA scores on the same test set — documented in prose on individual model pages, not as separate leaderboard columns).

Benchmark Specifications

FieldValue
Created byWMT (Conference on Machine Translation), Kocmi et al.
Paper"Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets" (ACL Anthology 2025.wmt-1.22)
Translation directions12
Metric (tracked)XCOMET-XXL, 0-100 scale
Other metrics reportedCometKiwi (reference-free), GEMBA (LLM-judge)
SaturationLow — the 2025 edition was explicitly designed to resist saturation

How WMT25 Is Scored

Submitted translations for each of the 12 evaluated directions are scored against WMT25's human evaluation sets, using the same automatic-metric conventions as FLORES-200 (XCOMET-XXL as the primary reference-based metric Benchgen tracks). WMT's own official ranking also incorporates direct human evaluation, but LLM technical reports citing WMT25 typically report only the automatic-metric scores.

State-of-the-Art Results

RankModelScore (XCOMET-XXL)SourceDate
1Hy-MT2-7B63.86Hy-MT2 technical report2026-05
2Hy-MT2-30B-A3B62.89Hy-MT2 technical report2026-05
3Hy-MT2-1.8B50.30Hy-MT2 technical report2026-05

Scores sourced from the Hy-MT2 technical report's Table 2 (WMT25 column, XCOMET-XXL sub-metric). The paper also reports higher scores for large closed models (e.g. Gemini 3.1 Pro thinking-mode at 88.38) evaluated under the same test set — not yet added to Benchgen pending those models' own page/benchmark-table sync.

WMT25 on Benchgen

No Benchgen results yet — be the first to run WMT25.

WMT25 vs Other Benchmarks

BenchmarkWhat it testsTasksSaturation
WMT25General MT, deliberately hard/fresh test set12 directionsLow
FLORES-200General multilingual translation quality, 200 languages1,056 directionsMedium

Run WMT25 on Your Model

Benchgen can evaluate your model against WMT25's human-curated test sets with version-controlled, regression-tracked results — useful for catching translation-quality drift across model updates on the specific directions your product actually serves.

Frequently Asked Questions

What is WMT25? WMT25 is the 2025 edition of the Conference on Machine Translation's annual General Machine Translation shared task, evaluating systems on a freshly-curated, deliberately difficult test set across 12 translation directions.
What does a good score look like? On the XCOMET-XXL metric (0-100 scale), the strongest large models score in the low-90s; mid-sized open models commonly score in the 60s-80s, reflecting the harder, less-saturated test set compared to FLORES-200.
Who created WMT25? WMT25 was created by the Conference on Machine Translation (WMT) community, documented in Kocmi et al.'s 2025 findings paper.

Last updated 2026-08-31.