Microsoft reduces voice agent latency with three new MAI models

Microsoft AI is expanding its audio lineup with MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. Transcription starting within the first few hundred milliseconds, multilingual speech generation, and 45 seconds of audio generated with 150 ms latency directly target real-time voice agents.

Processing a sentence before it is finished

Applications no longer need to wait until someone finishes speaking to receive the first elements of a transcript. MAI-Transcribe-2-Streaming produces its first hypotheses just over 100 ms after receiving audio, then continuously revises them as more context arrives.

Those partial transcripts can be used before the final text is committed. A voice agent can therefore begin reasoning or call tools while the person is still speaking.

The model supports 60 languages with automatic continuous language detection. It currently ranks first on Artificial Analysis for accuracy across both final and partial transcripts. Microsoft also says its internal evaluations show words appearing twice as fast as its closest competitor for use cases such as live dictation and subtitling. The latter comparison comes from Microsoft's own testing. One voice across 23 languages

The other side of the lineup focuses on generation. MAI-Voice-2.1 supports 23 languages across 26 locales.

A single voice identity can move between languages without switching speakers. Microsoft says the model also adapts its accent to the language being spoken instead of carrying the accent associated with the original language.

Both new voice models support cloning from a few seconds of reference audio across all supported languages. Microsoft also describes built-in consent guardrails intended to prevent unauthorized use. 45 seconds of audio with 150 ms latency

MAI-Voice-2.1-Flash retains the language capabilities of MAI-Voice-2.1 while targeting high-volume, latency-sensitive applications.

Microsoft says Flash can generate 45 seconds of audio with 150 ms end-to-end latency. The company also reports 55% faster model inference and a roughly 60% lower cost than comparable models.

Pairing it with MAI-Transcribe-2-Streaming targets the complete voice-agent loop: listening, beginning to interpret the request, potentially using tools, and generating a response without accumulating excessive delays between each stage. From $0.54 per hour to $15 per million characters

MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year.

For speech generation, MAI-Voice-2.1 costs $22 per million characters, while MAI-Voice-2.1-Flash is priced at $15 per million characters.

Microsoft has also built Chatter, a demo combining the models inside the MAI Playground. All three are available through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live. MAI-Voice-2.1 and Flash are also offered through OpenRouter, with LiveKit support listed as coming soon.