Is AI Safe for Healthcare Translation

As AI translation increasingly becomes a part of all manner of systems, institutions in the healthcare field are long past general interest in the technology: They are either planning for it, trying it out, or using it already.

Much like it happened in the days before AI, some solutions are grown organically while others come via specialists, and not necessarily from the language industry. In fact, foundation large language models (LLMs), which are massive neural networks trained on vast amounts of data, are making it possible for many with access and know-how to create AI translation tools.

Proprietary LLMs are also part of the equation, and they might make more sense for some unique use cases, such as clinical trials for pharmaceuticals or medical devices. The bottom line is that AI translation is making inroads in the healthcare industry in many different ways.

Examples of medical, pharmaceutical, and healthcare companies implementing AI translation are no longer rare. What will take time to come by is more data about the results of implementing the technology, especially, an answer to the question of whether AI translation is safe in healthcare settings. Right now, the only answer is, “it depends.”

At the time of publication of this article, some AI tools are not performing at the level needed to be considered “safe” across the board. An October 2024 report by the Associated Press indicated that Whisper, OpenAI’s multilingual transcription tool, frequently hallucinates (fabricates content). The report states that a private company tool based on Whisper had already been used in about 7 million medical visits.

Model hallucination is just one example of what can go wrong with AI tools that make them not safe. But AI’s impact on providers, patients, insurance companies, and affiliated parties is of course two-fold, and much is to be considered when deciding on AI translation. 

On the one hand, there could be great benefits, including noticeably increased language access and efficiencies, potential new markets, and service enhancements. On the other hand, there could be issues that go beyond hallucinations, including propagated linguistic errors, IT issues, and data vulnerabilities. Institutions adding translation-as-a-feature (TaaF) to existing systems, or integrating translation technology from AI companies like OpenAI or language technology providers like DeepL, do well to approach AI from a risk management perspective.

Are Hybrid Models the Answer?

Limits in accuracy and cultural nuance, terminology issues, and other linguistic problems are what professional linguists excel at spotting and fixing. In the years since the launch of neural machine translation (MT), their expertise has been key to training translation engines and post-editing MT outputs. 

That same expertise is now critical to LLM training and fine-tuning, and AI translation correction. In a setting where results are expected in real-time, such as with AI interpreting (via speech-to-speech translation, S2ST), the expectation is that no human would intervene. However, to be “safe” such settings would necessarily imply a low risk for patients. 

Low-risk settings in healthcare could include the patient’s point of admission, administrative interactions for financial and insurance purposes, and potentially self-service triage stations where vitals and basic medical information can be entered by patients and processed by AI.

Likewise, AI could be sufficient and “safe” in certain controlled situations, such as an appointment within a clinical trial where information is being gathered for initial subject screening, for instance, in scripted interviews where they must provide equally controlled answers like “Yes” or “No.”

As the technology improves, a hybrid approach would imply that human linguists would continue correcting AI-generated transcripts and text translations, and interpreting in higher-risk situations such as any non-controlled, non-scripted medical interactions.

In the hands of professional linguists or even bilingual clinicians, things like instant AI terminology look up and AI agents for different quality assurance steps would also help create hybrid workflows. All potential solutions must nonetheless be accompanied by proper mitigation strategies (e.g., human oversight, improved training data, error detection algorithms, etc).

Ethics and Regulations

Umbrella AI legislation like the European Union AI Act and the United States Executive Order on AI are good starting points for any industry to start considering or refining standards of practice and ethics for AI use. But the creation of standards and guidelines will need to rely on shared data as AI translation is implemented.

The addition of patient protection mechanisms should help create a framework for an ethical and regulation-compliant AI in healthcare. This means that patients and medical providers should be involved in the formulation of such guidelines, as should language experts.

All sources consulted for this article represent those groups. And they show a consensus that AI has the potential to completely transform language services in healthcare, but that optimism should be accompanied by careful consideration, planning, and implementation.

Continuous research alongside responsible development and deployment should help mitigate risk for all involved, inform guidelines, and address and anticipate compliance issues.