| Rank | Model | Score |
|---|---|---|
| 1 | hy-mt2-7b | 63.86 |
| 2 | hy-mt2-30b-a3b | 62.89 |
| 3 | hy-mt2-1-8b | 50.3 |
1 phaseActive
WMT's annual General Machine Translation shared task — human-evaluation test sets, 12 translation directions. Metric: XCOMET-XXL (0-100).
Quick answer: WMT25 is the 2025 edition of the Conference on Machine Translation's (WMT) annual General Machine Translation shared task — a harder, human-curated test set explicitly designed to resist saturation ("Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets"), covering 12 translation directions. It's the academic community's flagship head-to-head MT evaluation, run fresh each year against new source text.
What it tests: Translation quality on freshly-curated, deliberately-difficult source text across 12 language directions, evaluated by both automatic MT metrics and human annotators. Why it matters: Unlike static benchmarks that saturate over time, WMT refreshes its test set every year specifically to stay hard — the 2025 edition's own title flags this as an explicit design goal. Known limitations: Only 12 directions per edition (versus FLORES-200's 1,056), and coverage skews toward higher-resource language pairs relevant to the shared task's participant pool.
The Conference on Machine Translation (WMT) has run an annual General Machine Translation shared task since 2006, each year assembling a fresh test set of untranslated source documents and collecting both automatic-metric and human-evaluation judgments of submitted systems' translations. The WMT25 edition explicitly responds to benchmark saturation in prior years by curating harder source material, aiming to keep the shared task discriminative between top-tier systems.
WMT25 evaluates across 12 translation directions per the shared task's own scope. As with FLORES-200, Benchgen tracks the reference-based XCOMET-XXL metric, since it's the metric most consistently reported by LLM technical reports citing WMT25 results (papers frequently additionally report CometKiwi and GEMBA scores on the same test set — documented in prose on individual model pages, not as separate leaderboard columns).
| Field | Value |
|---|---|
| Created by | WMT (Conference on Machine Translation), Kocmi et al. |
| Paper | "Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets" (ACL Anthology 2025.wmt-1.22) |
| Translation directions | 12 |
| Metric (tracked) | XCOMET-XXL, 0-100 scale |
| Other metrics reported | CometKiwi (reference-free), GEMBA (LLM-judge) |
| Saturation | Low — the 2025 edition was explicitly designed to resist saturation |
Submitted translations for each of the 12 evaluated directions are scored against WMT25's human evaluation sets, using the same automatic-metric conventions as FLORES-200 (XCOMET-XXL as the primary reference-based metric Benchgen tracks). WMT's own official ranking also incorporates direct human evaluation, but LLM technical reports citing WMT25 typically report only the automatic-metric scores.
| Rank | Model | Score (XCOMET-XXL) | Source | Date |
|---|---|---|---|---|
| 1 | Hy-MT2-7B | 63.86 | Hy-MT2 technical report | 2026-05 |
| 2 | Hy-MT2-30B-A3B | 62.89 | Hy-MT2 technical report | 2026-05 |
| 3 | Hy-MT2-1.8B | 50.30 | Hy-MT2 technical report | 2026-05 |
Scores sourced from the Hy-MT2 technical report's Table 2 (WMT25 column, XCOMET-XXL sub-metric). The paper also reports higher scores for large closed models (e.g. Gemini 3.1 Pro thinking-mode at 88.38) evaluated under the same test set — not yet added to Benchgen pending those models' own page/benchmark-table sync.
No Benchgen results yet — be the first to run WMT25.
| Benchmark | What it tests | Tasks | Saturation |
|---|---|---|---|
| WMT25 | General MT, deliberately hard/fresh test set | 12 directions | Low |
| FLORES-200 | General multilingual translation quality, 200 languages | 1,056 directions | Medium |
Benchgen can evaluate your model against WMT25's human-curated test sets with version-controlled, regression-tracked results — useful for catching translation-quality drift across model updates on the specific directions your product actually serves.
Last updated 2026-08-31.