Previous research has shown that automatic prompt optimization can improve LLM performance. However, most studies have compared optimized prompts against generic baselines and focused on tasks such as coding, classification, or question answering. Comparisons between automatically optimized prompts and prompts written by localization experts for translation-related tasks have been limited, the researchers highlighted.
To address this gap, Marina Sanchez-Torrón, Senior Linguistic Engineer at Smartling, together with Smartling Data Scientists Daria Akselrod and Jason Rauchwerk, compared prompts written by localization experts with prompts automatically optimized using AI across three localization tasks: terminology insertion, translation, and LQA.
Rather than relying on general benchmarks, the researchers evaluated the approaches using proprietary datasets annotated by professional linguists, making the evaluation more representative of production localization workflows.
The results suggest automatically optimized prompts can perform similarly to expert-written prompts for some tasks, although expert input continues to provide advantages in others.
For terminology insertion, automatically optimized prompts achieved glossary compliance rates comparable to prompts written by localization experts while improving terminology match rates for several models.
Translation results followed a similar pattern, with most differences between automatically optimized and expert-written prompts proving statistically insignificant.
Overall, the findings suggest automatically optimized prompts can improve simple baseline prompts and, in some cases, achieve results comparable to those developed by localization experts.
The results were different for language quality assessment, or LQA. Prompts written by localization experts performed better at detecting translation errors, while automatically optimized prompts were more effective at categorizing and describing errors once they had been identified.
The researchers conclude that no single approach is best. Instead, they argue that practitioners should choose the approach — or combination of approaches — that best fits their available resources, such as labeled data, human effort, computing resources, and the requirements of each task.
For LSIs and enterprises building AI-assisted localization workflows, the study addresses a practical question: how much of prompt development can be automated across different localization tasks?
The findings suggest automatically optimized prompts can narrow the gap with prompts written by localization experts for terminology insertion and translation. At the same time, they show that expert linguistic input remains particularly valuable for tasks requiring more nuanced quality judgments, such as LQA.
Looking ahead, the researchers suggest exploring hybrid approaches that combine expert-written prompts with automatic optimization. They also call for cost-benefit analyses comparing manual and automated prompt development across the product lifecycle.