“Please, read the preliminary results with a pinch of salt,” wrote Tom Kocmi, researcher at Cohere and WMT25 organizer, in a LinkedIn post.
WMT25 saw increased participation with 36 unique teams (43 registered, but 7 withdrew or were disqualified). The task evaluated 32 language pairs, mostly English→X; half will undergo human evaluation, while the rest rely solely on automatic ranking.
The evaluation combined LLM-as-a-Judge (GPT-4.1, CommandA), reference-based metrics (MetricX-24-Hybrid-XL and XCOMET-XL), and quality estimation (CometKiwi-XL). For low-resource languages like Bhojpuri and Maasai, chrF++ was used. This mix balances reference bias and ensures broader reliability.
Slator 2025 AI Dubbing Report
The 85-page report analyzes the supply and demand for AI dubbing and the technical and operational nuances in delivering AI dubbing across verticals.
Who’s on Top
In the preliminary automatic rankings, a system called Shy_hunyuan-MT ranked top in almost every language pair. The report provides no further details, with full system descriptions to appear in the forthcoming WMT25 Findings.
Kocmi told Slator that the system appears to have been trained to optimize directly for the metrics used in the evaluation — “using MT metrics as a reward model to train a large language model,” as he explained — an approach that can boost automatic scores but does not necessarily align with human judgments.
He added that this illustrates how automatic evaluation can favor systems tuned to specific benchmarks, noting that preliminary human assessments do not confirm the same lead.
Among general-purpose LLMs, Google’s Gemini-2.5-Pro and OpenAI’s GPT-4.1 were the strongest performers, consistently ranking in the top positions across multiple languages. Cohere’s CommandA followed closely, while DeepSeek-V3 (Experts Weigh In on DeepSeek AI Translation Quality) and Anthropic’s Claude-4 formed the second tier of strong but less consistent performers. Other major LLMs, including Meta’s Llama-4, Alibaba’s Qwen-3, and Mistral AI’s Mistral, generally ranked lower.
Commercial systems, like Google, Microsoft, or DeepL — anonymized as ONLINE-B, ONLINE-G, and ONLINE-W — delivered reliable but mid-ranking performance. They often outperformed smaller open-weight LLMs, but generally fell behind the latest frontier LLMs and optimized systems. See also Google Translate Almost Doubles its Language Coverage Overnight.
Check out our latest updates on AI Translation.