Muse Voice Transcribe: Streaming ASR, Diarization, and Adaptive Delay
Coffee Summary
- FACT: On 1 Sep 2026 Meta introduced Muse Voice Transcribe, a real-time audio model with streaming ASR, endpointing, and diarization for 20+ speakers.
- FACT: It processes 80 ms chunks and uses adaptive delay via RL on WER and delay rewards.
- FACT: Training covers 70+ languages, 25 verified, with code-switching and keyword/context biasing.
- CLAIM: Meta says it ranks first on streaming STT and public diarization benchmarks.
What happened
Muse Voice Transcribe combines streaming ASR, speaker diarization, and endpointing in one multimodal stack. Meta positions it for overlaps, interruptions, accents, bilingual switching, hour-plus audio, and 20+ speakers without required post-processing. Meta AI and Muse Code dictation can use it via Fn.
Why it matters
Voice agents, meeting copilots, and wearables need low latency, who-said-what, and end-of-speech detection. Most stacks bolt diarization and VAD onto separate ASR; Meta’s pitch is joint tokens on one streaming foundation.
What changed
Audio arrives in 80 ms chunks. The model emits text or continues listening, then flushes at stream end. Adaptive delay balances word error rate and time to final transcription. Speaker-turn and speech-endpoint tokens train jointly with ASR.
Limitations
Ranking claims are vendor-dated and not independently re-run. Validate overlap-heavy, code-switched audio on your own microphones.
AIImpish Take
Muse Voice Transcribe is Meta selling working ears: streaming ASR, speaker tags, and endpoints in one model. Treat rankings as claims, then benchmark your own audio.
AIImpish