A dynamic, comic-book style graphic featuring the text "VS" centered on a diagonally split background, with a yellow sunburst on the left and a red explosion on the right.

Large language models are widely used across localization workflows, from translation and terminology management to language quality assessment (LQA). Achieving consistent results depends not only on the selected model but also on the prompts that guide it. 

As Language Solutions Integrators (LSIs) and enterprises experiment with different models and prompting strategies, developing and maintaining those prompts has become a practical challenge. A prompt optimized for one model, task, or language pair may not perform as well after those variables change. Model updates can also make previously effective prompts less reliable.

A new study by researchers at Smartling, presented at the 26th Annual Conference of the European Association for Machine Translation (EAMT 2026), held June 15-18 in Tilburg, the Netherlands explores whether part of that work can be automated.

Previous research has shown that automatic prompt optimization can improve LLM performance. However, most studies have compared optimized prompts against generic baselines and focused on tasks such as coding, classification, or question answering. Comparisons between automatically optimized prompts and prompts written by localization experts for translation-related tasks have been limited, the researchers highlighted.

To address this gap, Marina Sanchez-Torrón, Senior Linguistic Engineer at Smartling, together with Smartling Data Scientists Daria Akselrod and Jason Rauchwerk, compared prompts written by localization experts with prompts automatically optimized using AI across three localization tasks: terminology insertion, translation, and LQA.

Rather than relying on general benchmarks, the researchers evaluated the approaches using proprietary datasets annotated by professional linguists, making the evaluation more representative of production localization workflows.

The results suggest automatically optimized prompts can perform similarly to expert-written prompts for some tasks, although expert input continues to provide advantages in others.

Automation Performs Well on Translation and Terminology

For terminology insertion, automatically optimized prompts achieved glossary compliance rates comparable to prompts written by localization experts while improving terminology match rates for several models. 

Translation results followed a similar pattern, with most differences between automatically optimized and expert-written prompts proving statistically insignificant.

Overall, the findings suggest automatically optimized prompts can improve simple baseline prompts and, in some cases, achieve results comparable to those developed by localization experts.

Expert-Written Prompts Perform Better on LQA

The results were different for language quality assessment, or LQA. Prompts written by localization experts performed better at detecting translation errors, while automatically optimized prompts were more effective at categorizing and describing errors once they had been identified.

The researchers conclude that no single approach is best. Instead, they argue that practitioners should choose the approach — or combination of approaches — that best fits their available resources, such as labeled data, human effort, computing resources, and the requirements of each task.

For LSIs and enterprises building AI-assisted localization workflows, the study addresses a practical question: how much of prompt development can be automated across different localization tasks?

The findings suggest automatically optimized prompts can narrow the gap with prompts written by localization experts for terminology insertion and translation. At the same time, they show that expert linguistic input remains particularly valuable for tasks requiring more nuanced quality judgments, such as LQA.

Looking ahead, the researchers suggest exploring hybrid approaches that combine expert-written prompts with automatic optimization. They also call for cost-benefit analyses comparing manual and automated prompt development across the product lifecycle.