Back
Microsoft launches MAI-Transcribe-2-Streaming and two new voice models for real-time voice agents
SiTech AI Team3 min read

Microsoft launches MAI-Transcribe-2-Streaming and two new voice models for real-time voice agents

Microsoft released three new AI models: MAI-Transcribe-2-Streaming for real-time transcription and two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The streaming model ranks first on Artificial Analysis for transcription accuracy.

Microsoft released three new AI models on October 1: a real-time transcription model, MAI-Transcribe-2-Streaming, and two speech-synthesis models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.

Streaming transcription: 60 languages and record accuracy

MAI-Transcribe-2-Streaming is the company's first streaming transcription model: instead of waiting for a speaker to finish, it ingests continuous audio and returns text in stages. It supports 60 languages, detects them automatically, and ranks first on Artificial Analysis for final and partial transcript accuracy. The September 28 leaderboard shows a 2.5% word error rate for final transcripts, 2.8% for the first partial, and 0.13 seconds to final transcription.

Microsoft says first partial hypotheses arrive in just over 100 milliseconds, with words appearing as early as 320 milliseconds in most cases. For dictation and subtitling, internal evaluations show words appearing twice as fast as with the closest competitor. Introductory pricing is $0.54 per hour of audio through the end of the year.

Two new voice models

MAI-Voice-2.1 is Microsoft's strongest multilingual text-to-speech model, supporting 23 languages and 26 locales; a single voice can speak all of them with a native accent. It is priced at $22 per 1 million characters. MAI-Voice-2.1-Flash supports the same languages but is built for high-volume, latency-sensitive work: 45 seconds per generation, 150 milliseconds end-to-end, 55% faster inference and roughly 60% cheaper than comparable models, at $15 per 1 million characters.

Both models can clone any voice from a few seconds of reference audio, with built-in consent guardrails. In a 4,000-listener Turing test, 50.3% rated MAI-Voice as equally or more human-like than human recordings.

Why it matters

A voice agent runs on a loop: it has to hear, understand, decide and speak within the window in which a conversation still feels natural. Microsoft says pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash buys back time on both ends of that loop. Microsoft pitches them at contact centers, multilingual assistants and interactive learning.

RuntimeWire's analysis notes the $0.54 rate is 5.4 times the $0.10 per audio hour of September's batch-oriented MAI-Transcribe-2, though the workloads differ and the earlier price was a limited-time offer. It also cautions the latency figures describe reported model performance, not guaranteed end-to-end response times. Mustafa Suleyman, who leads Microsoft AI, said on X the launch is the world's most accurate real-time transcription model and is 55% faster and 60% cheaper than ElevenLabs. All three models are available through Microsoft Foundry, the MAI Playground, Vercel and Azure Voice Live, with the voice models also on OpenRouter and LiveKit coming soon.

Sources: TipRanks · Techmeme

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.