On October 1, Microsoft AI shipped its first streaming transcription model. MAI-Transcribe-2-Streaming. Sixty languages, continuous language detection, number one on Artificial Analysis partial and final streaming transcript accuracy. The core claim: latency sits on the Pareto frontier too, not just accuracy.
Paired with MAI-Voice-2.1 and Flash, Speech to Agent to Speech can run on Microsoft stack alone.
What shipped: a listen-and-speak set
| Model | Detail |
|---|---|
| Transcribe-2-Streaming | Real-time transcription, first partials just past 100ms after audio, 60 languages |
| Voice-2.1 | 23 languages and 26 locales, one voice with native accents, clone from seconds of reference |
| Voice-2.1-Flash | 45 seconds of audio at 150ms end-to-end, $15 per million chars (vs $22) |
| Price | Transcribe at $0.54 per audio hour (introductory through year-end) |
The internal-eval line amuses. Words land in transcripts twice as fast as the nearest rival. Partials arrive mid-sentence, so voice agents start reasoning and tool calls before sentences end.
Before: speech ends → full recognition → reason → speak
Now: mid-speech → partials → early reason and tools → speak
→ "working while listening" arrives
The Gemini Live post called in-conversation tool use the core. Partials are its start button.
Why 4 latencies: one number lies
The brief's action unfolds into a scorecard.
4 voice-agent latencies:
1. speech to partial: speech becoming partials (target the 100ms band)
2. partial to final: until the confirmed transcript
3. final to first token: agent reasoning to first token
4. first token to first audio: TTS until first sound
One total hides which stage lags
→ measure 4 apart and bottlenecks show
It applies the first-token versus completion split from the speed guide to voice. Voice adds TTS (Flash-class 150ms) on stages 3 and 4, so slices run finer.
Design fallout: interrupts and caching
Three rules for the partial era.
1. Interrupts as first class:
- must cut in before speech ends
- revise reasoning when partials revise (treat pre-final as hypotheses)
2. Separate caching:
- cache before search and STT on repeats ([caching post](/en/blog/prompt-caching-ai-api-cost/))
- voice repeats more than text (greetings, checks, corrections)
3. Accuracy with latency:
- score partial and final accuracy apart, AA style
- check Pareto membership (accurate but slow loses in voice)
It links to the multimodal pipeline from the Qwen Omni post. Listening, seeing, and speaking on one pipeline means splitting the latency budget.
CodeBridge Mini Lab: time 4 latencies
1. Fix 10 voice questions (5 short orders plus 5 long briefs)
2. Time 4 stages (stopwatch, logs):
[ ] speech to partial (target the 100ms band)
[ ] partial to final
[ ] final to first token
[ ] first token to first audio
3. Call it:
- stage 1 lags → review STT (compare Transcribe-class)
- stage 3 lags → retune reasoner and effort ([Sonnet high](/en/blog/sonnet-5-5-high-effort-sweet-spot/))
- stage 4 lags → review TTS (compare Flash-class)
- all small → lift feel with interrupt UX
It ports the 500-token feel timing from the speed guide to voice.
Conclusion: listening speed is agent speed
One line to close.
Voice agent UX rides on STT partials, TTS, and interrupts as much as LLM reasoning.
The MAI set shows listen-and-speak standardizing into parts. Transcribe listens, Voice speaks, your agent fills between. One task for today: time the 4 latencies of your voice pipeline. Slow stages named tell you what to change.
Further reading
- Voice AI through Gemini 3.8 Live
- AI model speed and latency guide
- Multimodal AI through Qwen3.8 Omni Flash
References
- Microsoft AI: Our first streaming transcription model (Oct 1, 2026)
- Microsoft Learn: MAI-Transcribe-2-Streaming overview
- The Decoder: Microsoft AI releases transcription and TTS models
Go deeper with a course
To design listen, judge, and speak pipelines as a structure, this course builds harness, loop, and graph layers exactly like the 4 latencies here.