A screenshot of the LALAL.AI website homepage on a yellow background, featuring the headline "Remove Vocals and Instrumentals from Audio and Video" alongside a file upload box with a "Select Files" button.

The first LALAL.AI model built exclusively for voice isolation gives localization teams and dubbing SaaS providers cleaner dialogue tracks, faster pipeline throughput, and a dedicated API for batch processing at scale. 

Zug, Switzerland, July 14LALAL.AI today announced Lynx, its first neural network built from the ground up exclusively for voice isolation. The model is now available through the LALAL.AI API, web, desktop (cloud mode), and mobile apps and powers the company’s Voice Cleaner and Voice & Noise stem, the tools most commonly used by localization companies, dubbing SaaS platforms, and post-production teams that need clean dialogue as the foundation of their workflows.

Lynx is trained to separate speech from everything else: background music, crowd noise, mechanical interference, environmental sounds, and the full range of acoustic artifacts present in content produced outside controlled studio conditions. The result is a clean voice track ready for transcription, localization, text-to-speech input, or human voice replacement, without the manual cleanup steps that slow batch production.

“It’s easier to describe what Lynx leaves untouched (the voice itself) than to list every category of noise it can handle,” says Nik Pogorsky, LALAL.AI Product Owner and Co-founder. “We consider the resulting training set one of the cleanest we’ve used. It totals several hundred hours of carefully filtered audio, supplemented by several thousand additional hours of more loosely cleaned content used during pretraining.”

Lynx runs on a proprietary architecture developed over one year by the LALAL.AI research team. The model is six times smaller than Andromeda, LALAL.AI‘s flagship cloud stem separation model, which reduces computational load without compromising output quality. This is a meaningful factor for SaaS providers and localization platforms processing high volumes of content through the API. The model was trained on a manually curated dataset: since no reliable automated pipeline exists for cleaning the full dynamic range of real-world speech from diverse noise sources, the team spent months hand-selecting, filtering, trimming down, and preparing audio tracks covering conditions from quiet interviews to noisy field recordings.

The demand for dedicated voice cleaning in localization workflows is reflected in LALAL.AI‘s own usage data: the Voice & Noise stem, now powered by Lynx, is the platform’s second most-used separation track, with more than 11.5 million audio splits processed on it in 2025. Dubbing and localization providers are among its primary professional users, including multilingual platforms supporting enterprise clients across media, e-learning, and corporate communications, and AI-powered dubbing SaaS products that have integrated the LALAL.AI API directly into their processing pipelines as a core audio preparation step.

Lynx is available now through the LALAL.AI API for B2B integrations, SaaS platforms, and enterprise deployments, as well as through the Voice Cleaner product via browser, mobile app, and desktop app cloud mode. Planned near-term improvements to the Lynx architecture include enhanced separation of choral singing and lead vocals, and improved isolation of speech recorded at distance from the microphone, both relevant to content types commonly handled in localization pipelines.

About LALAL.AI

LALAL.AI is an AI-powered audio processing platform offering stem separation, voice cleaning, and noise removal tools for creators, developers, and enterprises. The platform is accessible via web, mobile, desktop application, API, and VST plugin.