Qwen3.8-LiveTranslate separates voices during translation
Qwen3.8-LiveTranslate traduit 60 langues, sépare les interlocuteurs en temps réel et réduit son décalage moyen à 2,3 secondes, avec restitution vocale dans 29 langues.
A bilingual meeting does not become easier to follow when every contribution appears to come from the same person. With Qwen3.8-LiveTranslate, Qwen is attempting to translate the conversation while separating speakers, preserving their vocal characteristics, and displaying the original remarks at the same time.
Qwen3.8-LiveTranslate succeeds version 3.5 with an architecture designed to process the audio it hears, recognized text, translation, and generated speech within a single stream. The team says this structure improves faithfulness, fluency, and concision while reducing the delay accumulated during interpretation.
The average lag reported by Qwen falls from 2.8 to 2.3 seconds. This figure refers to LAAL, a metric used to assess how far simultaneous translation trails behind the source speech. It should not be confused with the exact delay experienced by every user.
Network quality, distance from the data center, microphone capture, buffering, and spoken output can each introduce additional latency. The 2.3-second figure therefore describes Qwen’s published test result, not a constant guarantee for every conversation.
The main change concerns how processing is organized. Qwen describes an Interleave architecture in which audio, source and translated text, and generated speech are woven together according to their temporal order.
The system does not necessarily wait for a speaker to finish before beginning the translation. It processes the stream as it arrives, reuses what it has already heard, and considers translation fragments it previously generated.
This continuity is intended to prevent a sentence from being treated as a series of isolated elements. When translating between languages with different word orders, the system may need to wait for information that arrives later before phrasing the beginning correctly. It may also anticipate a unit of meaning, then revise it as the context becomes clearer.
Qwen3.8-LiveTranslate uses a Hybrid-MoE structure divided between two modules called Thinker and Talker. The first receives the information required for comprehension and translation. The second turns the translated text into speech.
Thinker places audio, optional images, the source transcript, and the translation into a single time-ordered sequence. Speech understanding and text generation are therefore not described as two fully separate services connected only after processing.
Talker then receives the translation along with information from the original audio. It produces the corresponding speech while attempting to preserve the original speaker’s vocal characteristics.
This preservation does not reproduce every aspect of a voice identically. Prosody, accent, emotion, and pacing may change when moving from one language to another. Some constructions also require a longer or shorter sentence than the original.
Speaker separation takes place during the conversation. When several people speak in turn, the system assigns their contributions to distinct identifiers and maintains that distinction in both the text and spoken output.
Qwen’s phrase “Names the speaker” can be misleading. The model separates and labels voices, but it does not automatically determine anyone’s real-world identity. An identifier such as “Speaker 1” still needs to be mapped to a name by the application or user when that information is required.
This real-time diarization should make meetings, interviews, and other multi-speaker exchanges easier to follow. It may also prevent the cloned voice from changing character whenever another person begins speaking.
The demonstration primarily covers people speaking one after another. Qwen has not published detailed measurements for overlapping speech, extremely noisy settings, or speakers with similar vocal characteristics.
The system produces the source transcript and translation in parallel. An interface can display both versions in sync, allowing users to check a term, name, or phrasing without returning to a separate transcription service.
This bilingual output extends beyond display. It can provide a foundation for subtitles, meeting organization, content indexing, or later retrieval of a specific passage.
Conversational history also helps resolve ambiguity. A proper name that was misheard at the beginning of a discussion may become easier to identify after it is repeated, connected to a company, or accompanied by visual information.
The model draws on previous turns to maintain more consistent terminology. An abbreviation, reference, or word with several meanings can be translated according to what has already been said rather than from the current sentence alone.
Priority terms can also be registered in a session corpus. The API accepts up to 1,000 mappings intended to improve the translation of brand names, professional expressions, and specialized terminology.
Visual context supplements the audio. An application can send an image or frames from a video stream to help the model interpret a gesture, displayed text, an object, or the surrounding situation.
This feature does not mean Qwen3.8-LiveTranslate generates or edits video. Audio remains mandatory, while images are optional inputs used to reduce certain ambiguities.
Language coverage reaches 60 languages for audio input and text translation. These include French, English, Chinese, Spanish, German, Italian, Portuguese, Japanese, Korean, Arabic, Russian, Hindi, Turkish, and Vietnamese.
Spoken output is limited to 29 languages. The remaining 31 can be understood and translated as text, but cannot necessarily be spoken by the service. The 60-language figure should therefore not be interpreted as 60 complete voice outputs.
Qwen evaluates speaker separation on Omnilingua-MSpeaker, a dataset for long recordings involving multiple speakers across 14 translation directions. The company claims that its model outperforms several competing systems on faithfulness, fluency, concision, and diarization error rate.
A second evaluation uses the