LAUNCH

Meta’s Muse tops streaming speech-to-text with 3.1% error

Signals Inbox·September 2, 2026·Voice AI

Meta introduced Muse Voice Transcribe, a real-time speech model that handles transcription, speaker diarization and endpoint detection inside one system. It ranks first on current streaming speech benchmarks and supports more than 25 languages, pushing voice AI closer to reliable always-on perception rather than a batch transcription feature.

The Signal, Explained in 3 Minutes

Q1What did Meta officially launch?

Meta AI Research's official post introduces Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs.

Q2What does it do in one model?

It performs streaming automatic speech recognition, diarization for more than 20 speakers and endpoint detection, while supporting more than 25 languages and code-switching.

Q3How good is it?

Meta says Muse ranks first on Artificial Analysis for streaming speech-to-text and on public diarization benchmarks as of September 1. The title's 3.1% error figure reflects benchmark performance rather than every real-world environment.

Q4Why combine diarization and transcription?

Voice agents need to know not only what was said but who said it and when a turn ended. Combining those functions reduces latency and complexity versus chaining several separate models.

Q5What becomes possible?

More reliable live meeting agents, call-center systems and conversational assistants that listen continuously and react in real time. Speech recognition is becoming a perception layer rather than a transcription utility.

← Back to the signals