WMT 25 Results

Like every year, the WMT General Machine Translation Shared Task, one of the most closely watched scientific benchmarks, offers an early snapshot of the evolving AI translation landscape. For details, see WMT24 Pits LLMs Against NMT in Preliminary Ranking Based on Automatic Metrics.

On August 11, 2025, the organizers released the WMT25 preliminary rankings of participating systems based on automatic evaluation.

Caveat: while useful for identifying trends, the rankings are not final. The official results will be published in November after human evaluation, which the organizers stress is “more reliable” and will “supersede the automatic results.”

“Please, read the preliminary results with a pinch of salt,” wrote Tom Kocmi, researcher at Cohere and WMT25 organizer, in a LinkedIn post.

WMT25 saw increased participation with 36 unique teams (43 registered, but 7 withdrew or were disqualified). The task evaluated 32 language pairs, mostly English→X; half will undergo human evaluation, while the rest rely solely on automatic ranking.

The evaluation combined LLM-as-a-Judge (GPT-4.1, CommandA), reference-based metrics (MetricX-24-Hybrid-XL and XCOMET-XL), and quality estimation (CometKiwi-XL). For low-resource languages like Bhojpuri and Maasai, chrF++ was used. This mix balances reference bias and ensures broader reliability.

Who’s on Top

In the preliminary automatic rankings, a system called Shy_hunyuan-MT ranked top in almost every language pair. The report provides no further details, with full system descriptions to appear in the forthcoming WMT25 Findings.

Kocmi told Slator that the system appears to have been trained to optimize directly for the metrics used in the evaluation — “using MT metrics as a reward model to train a large language model,” as he explained — an approach that can boost automatic scores but does not necessarily align with human judgments. 

He added that this illustrates how automatic evaluation can favor systems tuned to specific benchmarks, noting that preliminary human assessments do not confirm the same lead.

Among general-purpose LLMs, Google’s Gemini-2.5-Pro and OpenAI’s GPT-4.1 were the strongest performers, consistently ranking in the top positions across multiple languages. Cohere’s CommandA followed closely, while DeepSeek-V3 (Experts Weigh In on DeepSeek AI Translation Quality) and Anthropic’s Claude-4 formed the second tier of strong but less consistent performers. Other major LLMs, including Meta’s Llama-4, Alibaba’s Qwen-3, and Mistral AI’s Mistral, generally ranked lower.

Commercial systems, like Google, Microsoft, or DeepL — anonymized as ONLINE-B, ONLINE-G, and ONLINE-W — delivered reliable but mid-ranking performance. They often outperformed smaller open-weight LLMs, but generally fell behind the latest frontier LLMs and optimized systems. See also Google Translate Almost Doubles its Language Coverage Overnight.

Check out our latest updates on AI Translation.