Home AI Meta Muse Voice Transcribe Streams Speech With 20+ Speakers and 25 Languages

Meta Muse Voice Transcribe Streams Speech With 20+ Speakers and 25 Languages

Meta Superintelligence Labs' first real-time audio model pairs streaming ASR with native diarization and endpointing.

0
Image: Meta.

Meta’s newest ears can follow a crowded room. The Muse Voice Transcribe model streams speech to text in real time. It tags who said what across more than 20 speakers, and it follows sentences that jump between languages. Meta Superintelligence Labs shipped it on September 1. Engadget flagged the launch, but the substance below comes from Meta’s own research blog.

One model, three jobs

Most pipelines split transcription, speaker tagging and endpoint detection into separate tools. Muse Voice Transcribe, instead, handles all three natively. It is an autoregressive multimodal model from the Muse Spark family. Audio arrives in 80 millisecond chunks, and each chunk becomes one soft token. Then the model chooses: keep listening, or emit text. “The model decides when to listen.” Mark Zuckerberg explained that adaptive delay in his launch post, adding that it waits longer on hard words and commits faster on easy ones.

Because the model controls its own delay, reinforcement learning tunes the speed-accuracy trade-off directly. As a result, Meta says it reaches the Pareto front on time to final transcription.

Special tokens handle the extra jobs. A start-of-turn token flags a possible speaker change, while a speaker tag names who talked. For endpointing, speech-onset and speech-endpoint tokens bracket each utterance. Meta trained all three tasks together, and it stacked diarization and endpointing rewards on top of the ASR reward.

Meta frames the model as ears for personal superintelligence. In the blog’s roundtable transcript, researchers argue a personal agent must listen like a human in real conversations, not just wait for voice commands. Overlaps, interruptions and accents make that hard. Therefore streaming endpointing and diarization sit at the core of the design.

Benchmarks Meta claims

On Artificial Analysis, the model ranks first for streaming speech-to-text with a 3.1 percent final word error rate. It also tops public diarization benchmarks at 17.5 percent error. Moreover, Meta’s adaptive-delay chart places it below the previous Pareto frontier, near 3.0 percent error at 0.16 seconds. Rivals on that chart include Soniox, Cartesia, ElevenLabs and Gemini 3.5 Transcribe Live. Still, these are self-reported numbers, so independent runs will matter.

Meta chart showing Muse Voice Transcribe below the previous Pareto frontier on streaming word error rate versus time to final transcription
Meta’s chart puts Muse Voice Transcribe ahead of Soniox, Cartesia and Gemini on the speed-accuracy curve. Image: Meta.

Languages, long sessions, and biasing

Training covered more than 70 languages, yet Meta validated 25 at launch and recommends starting there. Code-switching works natively, both inside a sentence and between sentences. The blog demo even transcribes a Mandarin-English mix about a doctor appointment. Long audio is no problem either: sessions beyond one hour with 20 or more speakers run without post-processing. Finally, language, keyword and context biasing let the model know your world. Say “Meta,” “Muse” or “Menlo Park,” and it spells them right.

Where you can try it

The model already powers dictation inside the Meta AI Mac app, and that engine feeds voice features in other apps too. Developers get it through Muse Code and the Meta Model API today. Pricing runs about $3 per 1,000 audio minutes, or 18 cents an hour on Meta’s rate card. Streaming and file transcription cost the same. However, rate limits cap users at eight concurrent streams and 1,000 streams per hour. A live demo also runs on the research blog.

The race with Google

Meta’s launch lands less than a week after Google’s Gemini 3.5 Transcribe. Google is baking its model into Android and, eventually, Chrome. Meanwhile Meta has not said whether Muse Voice Transcribe will reach its flagship apps. So the fight is platform versus API, for now.

The model is also MSL’s latest in a busy month. The lab recently shipped the Muse Code coding agent, the open-weight Muse Glimmer model and the Meta AI Mac app. Voice is the new keyboard, after all. From Google’s on-device dictation to Claude’s voice mode, every major lab now ships speech-first interfaces. Code-switching support, which handles mid-sentence language jumps, is Meta’s party trick. Whether the benchmarks hold outside Meta’s own rooms will decide how loud the applause gets.

NO COMMENTS

Exit mobile version