Respondents also expressed difficulty resulting from inadequate translation-as-a-feature (TAAF, i.e., translation incorporated into other platforms like content management systems). Specifically, TAAF quality is not seen as good enough to forgo human intervention, including post-editing. Respondents report “friction” when explaining the need for human quality reviews to stakeholders.
In general, respondents emphasized the need for clarity about the current limitations of AI tools, and the importance of “dispelling misconceptions about AI as an all-in-one solution.”
Users of AI language solutions should understand that not all large language models (LLMs) are created equal. They are trained using different datasets, and developers often have specific goals and features in mind when creating a new tool powered by an LLM.
Some issues that often still require human intervention (depending on the particular use case and AI translation tool) include AI translation outputs that lack natural-sounding style and voice, inaccurate conversion of domain-specific content, and weaker performance with low-resource languages, among others.
It is also important to recognize that even models released from the same developer may produce noticeably different translation quality, as they are likely to have different features and training.
This was apparent in Slator’s previous limited tests using OpenAI’s broad multimodal model GPT-4o, and GPT-o1, a model developed for better reasoning ability, which revealed some AI translation issues.
When Slator asked GPT-4o to translate a sample text from English to French, Spanish, and German, the translations tended to have high general accuracy but were weaker in terms of natural style. The model was able to handle domain-specific terminology, acronyms, and slang, yet offered translations so literal that “a native speaker would immediately identify [the translations] as unedited machine translation (MT).”
Conversely, when Slator informally tested the translation ability of ChatGPT-o1 (released in December 2024) the model produced generally accurate translations that seemed more natural, and stylistically nuanced. However, the model struggled with some acronyms, failing to transliterate them into non-Latin scripts, and struggled with right-to-left oriented scripts, particularly in terms of punctuation.
2024 Slator Pro Guide: Translation AI
The 2024 Slator Pro Guide presents 20 new and impactful ways that LLMs can be used to enhance translation workflows.
High-Risk Domain Considerations
When considering AI translation solutions, a domain-specific risk analysis is always warranted. Users who do not speak the AI-translated/generated language will not be able to tell whether the final quality is acceptable for their needs without calling on expert reviewers.
For non-language-specific purposes, and despite continued improvements, AI hallucinations (fabricated responses to queries) are a persistent concern. As recently as October 2024, the Associated Press (AP) reported an example of significant, and frequent, hallucinations from OpenAI’s multilingual AI transcription model Whisper. In some examples, the model even added racial or violent commentary to transcripts that were absent from recorded audio.
This is cause for concern as, according to the AP article, a Whisper-based tool was already being used in several medical institutions, risking inaccurately recorded communication between providers and patients.
The risk of poor performance from AI tools in such high-stakes use cases highlights the need, as one respondent in the Slator survey stated, “to align company expectations around AI and software in general (reducing cost and TAT) with a robust localization process.”
By fostering a more realistic understanding of what AI translation tools can achieve without human intervention, organizations can navigate the evolving language technology landscape more effectively, and set appropriate expectations for quality, cost, and turnaround times.