Skip to content

Microsoft's new speech model transcribes in 0.13 seconds, tops streaming accuracy charts

· by Pondero Newsdesk

The short version

Microsoft shipped MAI-Transcribe-2-Streaming on October 1, 2026, plus two companion voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at real-time voice agents.

Microsoft's new speech model transcribes in 0.13 seconds, tops streaming accuracy charts

Microsoft's newest transcription model turns speech into a finished, stable transcript in roughly the time it takes to blink, and per Microsoft it still posts the lowest error rate on a leading streaming-accuracy benchmark.

What

Microsoft released MAI-Transcribe-2-Streaming on October 1, 2026, a real-time speech-to-text model that streams words as they're spoken instead of waiting for a recording to end. Per Microsoft's own announcement, the model holds the lowest final word error rate of any model on Artificial Analysis's streaming speech-to-text leaderboard, at 2.5%, while reaching that final transcript just 0.13 seconds after the speaker stops. That pairing sits on the benchmark's efficiency frontier, per Microsoft: rival models such as Deepgram Flux (7.4% WER, 0.02 seconds) and Cartesia Ink Preview (3.1% WER, 0.11 seconds) trade accuracy for speed or the reverse, while MAI-Transcribe-2-Streaming holds both. The model covers 60 languages with automatic, continuous language detection, so a conversation that switches languages mid-sentence doesn't need to be flagged in advance. Introductory pricing is $0.54 per audio-hour through the end of 2026, per Microsoft.

Two voice models shipped alongside it, detailed in the same announcement. MAI-Voice-2.1 generates speech in 23 languages across 26 locales and keeps one consistent voice as it switches languages, priced at $22 per million characters. MAI-Voice-2.1-Flash, built for high-volume workloads, generates up to 45 seconds of audio at an end-to-end latency of about 150 milliseconds, which Microsoft says is 55% faster and roughly 60% cheaper to run than comparable models, at $15 per million characters.

Voice agents can now react mid-sentence, not after it

The practical change is what a voice agent can do while someone is still talking. Microsoft says the model produces its first "partial" transcript hypotheses in just over 100 milliseconds of receiving audio, then revises them as more context arrives and locks a stable final version almost immediately after. That lets a tool call or an agent's next reasoning step start before the caller finishes a sentence, instead of waiting for a pause. For live captioning or dictation specifically, Microsoft's internal evaluations found words landing on screen twice as fast as its closest competitor. For teams building call-center bots, meeting-note tools, or live-subtitle features, the three models now form a matched set: transcribe what's heard, decide what to do, and speak the reply, all inside one vendor's latency budget.

Context and reactions

MAI-Transcribe-2-Streaming builds on MAI-Transcribe-2, the batch-only model Microsoft shipped September 3 with leading FLEURS and Artificial Analysis accuracy scores but no real-time option, per Microsoft's model page. The streaming variant adds the live feed Microsoft frames as closing the gap between batch accuracy and real-time usability rather than choosing one over the other. MAI-Voice-2.1 follows June's MAI-Voice-2, extending language coverage and adding the ability to hold one voice identity across languages rather than switching accents when the language changes. TestingCatalog first reported the launch on October 2, a day after Microsoft's own post went live.

What to watch next

Watch where the stack actually ships inside Microsoft's own products. Azure AI Foundry access is live now, but Copilot voice features and Teams real-time transcription are the bigger volume tests. Also worth tracking: whether third-party benchmarks replicate the 2.5% WER figure once developers outside Microsoft start running it against their own audio, and whether competitors on the Artificial Analysis board respond with their own latency-accuracy tradeoffs.

Sources