✉ The Friday AI Brief: the week's 5 best AI stories, tools & comparisons — in your inbox every Friday morning.

Microsoft logo

Microsoft Gives Voice Agents an Ear and a Voice: MAI-Transcribe-2-Streaming and Two New Speech Models

Microsoft Gives Voice Agents an Ear and a Voice: MAI-Transcribe-2-Streaming and Two New Speech Models

October 2, 2026 — Redmond

Microsoft AI has released its first streaming speech-to-text model, MAI-Transcribe-2-Streaming, alongside two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash — available in public preview through Microsoft Foundry. The takeaway: real-time conversations with AI agents get measurably closer to sounding like a phone call instead of a walkie-talkie.

What happened

Announced October 1 by Microsoft AI, the trio completes the voice side of Microsoft’s in-house model stack. MAI-Transcribe-2-Streaming transcribes live audio continuously as it arrives, returning partial hypotheses in the low hundreds of milliseconds and refining them as more context comes in. It covers 60 languages with automatic language detection, so speakers can switch languages mid-conversation without restarting the session. Microsoft says the model ranks first on Artificial Analysis benchmarks for both partial and final transcription accuracy — its own benchmark claim, not independent verification.

The transcription model is priced at $0.54 per audio hour on an introductory basis through the end of 2026, and it requires Azure Speech SDK 1.52.0. It currently runs from Sweden Central, Central US, and Southeast Asia, with East US 2 listed as coming soon.

The voice generation side

MAI-Voice-2.1 is built for expressive, multilingual speech: 23 languages across 26 locales with a single consistent cross-language voice, priced at $22 per million characters. MAI-Voice-2.1-Flash trades some expressiveness for speed and cost — 150 milliseconds of latency at $15 per million characters. Both can clone a voice from a few seconds of reference audio, though Microsoft says voice cloning requires access approval and consent safeguards. A demo voice agent called “Chatter” is available on the MAI Playground, and the voice models are also listed on OpenRouter.

Why it matters

Voice is becoming one of the most important interfaces for AI agents, and latency can matter as much as accuracy: an agent that takes several seconds to respond can feel unusable no matter how good its language model is. Combined, the three models let developers build a full conversational loop — hear, think, speak — on Microsoft’s own stack, putting Microsoft in direct competition with voice offerings from OpenAI and Google. The standard caveat applies: everything is in public preview with no service-level agreement, so don’t run production workloads on it yet.

Frequently asked questions

What is MAI-Transcribe-2-Streaming?
Microsoft’s first streaming speech-to-text model. It transcribes live audio continuously and returns partial results in the low hundreds of milliseconds, instead of waiting for a speaker to finish.

How much does it cost?
The streaming transcription model costs $0.54 per audio hour through the end of 2026 (introductory pricing). MAI-Voice-2.1 is $22 per million characters; the Flash variant is $15 per million.

How many languages are supported?
The transcription model supports 60 languages with automatic detection. The voice models support 23 languages across 26 locales.

Can MAI-Voice clone a voice?
Yes — from a few seconds of reference audio. Microsoft says cloning requires access approval and consent safeguards to prevent misuse.

Where can I try the models?
They’re in public preview via Microsoft Foundry, with a demo agent called “Chatter” on the MAI Playground. The voice models are also on OpenRouter.

Sources: Microsoft AI; TradingView / Seeking Alpha; The Decoder; WindowsReport; RobotToday.

Leave a Comment

Your email address will not be published. Required fields are marked *

Get the 5 best AI tools every week

Top AI news, tools, and prompts — one short email. Free, unsubscribe anytime.

Scroll to Top