Skip to main content
AI Interview Question
INTERVIEW GUIDEAI News7 questions5 min readOct 3, 2026

MAI-Transcribe-2-Streaming: partials, latency budgets, and preview risk in interviews

Microsoft shipped MAI-Transcribe-2-Streaming on 1 Oct 2026: partials in just over 100 ms, paired voice models, and a public-preview caveat for interviews.

MAI-Transcribe-2-Streaming: partials, latency budgets, and preview risk in interviews

On 1 October 2026, Microsoft AI announced MAI-Transcribe-2-Streaming, its first streaming transcription model, together with MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The company frames the three as building blocks for low-latency conversational voice agents. The announcement post is at https://microsoft.ai/news/our-first-streaming-transcription-model/

MAI-Transcribe-2-Streaming does real-time speech-to-text in 60 languages with automatic continuous language detection. Microsoft says first partial hypotheses appear in just over 100 ms after audio arrives, then revise as context grows before a stable final transcript is committed. That design is meant so an agent can start reasoning or tool calls before the speaker finishes.

Microsoft states the model ranks no. 1 on Artificial Analysis for accuracy of both final and partial transcripts and sits on the accuracy-versus-latency Pareto frontier. Its model card, updated 1 October 2026, cites an Artificial Analysis final-transcription average WER of 2.5% for MAI-Transcribe-2-Streaming. Those leaderboard figures are a snapshot with Artificial Analysis methodology. They are not a universal production guarantee. Re-check live rankings and run your own evals. The model card is at https://microsoft.ai/pdf/MAI-Transcribe-2-Streaming-Model-Card-Memo.pdf and the streaming STT leaderboard is at https://artificialanalysis.ai/speech-to-text/streaming

Introductory pricing for MAI-Transcribe-2-Streaming is $0.54 per hour of audio through 31 December 2026. MAI-Voice-2.1 is priced at $22 per 1M characters. MAI-Voice-2.1-Flash is priced at $15 per 1M characters. Microsoft claims Flash has about 55% faster inference and about 60% cheaper pricing than comparable models, plus about 150 ms end-to-end latency to generate 45 seconds of audio. Those speed, cost, and latency figures are Microsoft claims. These sources do not state p99 SLOs beyond the first-partial and Flash figures Microsoft published.

All three models are available through Microsoft Foundry. The blog also lists MAI Playground, Vercel, and Azure Voice Live, with LiveKit coming soon. The voice models are also listed via OpenRouter. Integration paths for streaming STT include an OpenAI Realtime-compatible WebSocket API and the Azure Speech SDK. Catalog entry: https://ai.azure.com/catalog/models/MAI-Transcribe-2-Streaming?publisher=microsoft

Microsoft Learn documents MAI-Transcribe-2-Streaming as public preview, without an SLA, and not recommended for production workloads. That Learn note is at https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming Do not conflate this streaming model with the separate batch MAI-Transcribe-2 listing.

In an interview, a voice agent is a closed latency loop: hear, decide or call tools, then speak. Ask whether the system acts on partial transcripts mid-utterance or waits for finals, how you budget time for that loop, and how you would gate a streaming STT swap while the API is still preview with no SLA. A vendor leaderboard win is useful context. It is not enough without your own evals and a clear preview risk plan.

https://microsoft.ai/news/our-first-streaming-transcription-model/ https://microsoft.ai/pdf/MAI-Transcribe-2-Streaming-Model-Card-Memo.pdf https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming https://artificialanalysis.ai/speech-to-text/streaming https://ai.azure.com/catalog/models/MAI-Transcribe-2-Streaming?publisher=microsoft

["Microsoft AI""Speech-to-text""Voice agents""Streaming""Latency""MAI-Transcribe""Public preview"]

Questions in this guide

Deep explanations with architecture diagrams for every question below.