Meta's Superintelligence Labs has released Muse Voice Transcribe, its first model built for real-time audio. It returns text while someone is still talking, separates who is speaking, and decides where an utterance ends, all inside a single model. Pricing starts at 0.18 USD (about 28 yen) per hour, less than half of what comparable services charge.
The price of 0.18 USD an hour
Real-time transcription has long been expensive relative to what it does. Meta's figure is 3 USD (about 460 yen) per 1,000 audio minutes, which works out to 0.18 USD (about 28 yen) per hour.
Lined up against rivals, the gap is clear. Cartesia's Ink-2 runs 4 USD (about 610 yen) per 1,000 minutes, while ElevenLabs Scribe v2 Realtime and Deepgram Flux both sit at 6.5 USD (about 990 yen). Meta landed 25 percent below even the cheapest of them.
※1 USD = 153 JPY (as of September 9, 2026)
This corner of the market was already crowded. OpenAI shipped GPT-Realtime-Whisper for the same job in May and cut transcription prices again in July. Rather than fight on peak performance, Meta is undercutting the field, the same approach it took with Muse Spark 1.1 and 1.2.
Deciding every 80 milliseconds whether to wait or commit
The technical core is how the model handles delay. Muse Voice Transcribe slices incoming audio into 80-millisecond chunks, and after each one it decides whether to keep listening or emit the next word as text.
Waiting longer gathers more context and improves accuracy, but pushes output further behind the speaker. Rushing raises the error rate. The model varies that wait per word, depending on how hard the word is: easy ones come out quickly, ambiguous ones get more listening time.
Meta trained this behavior with reinforcement learning, a method that shapes a policy by rewarding desired outcomes. Rewarding both a low error rate and short delay at once keeps the model from collapsing into either extreme.
Speaker separation and endpointing in the same model
Real-time audio pipelines have typically used separate systems for transcription, speaker separation, and end-of-speech detection. Muse Voice Transcribe folds all three into one model.
For speakers, it writes turn changes directly into the running text and tags each passage with an identifier from A to Z. It can distinguish more than 20 speakers at once and handle recordings longer than an hour without post-processing. In a demo with eight people in one room, the system assigned utterances to individuals as they happened.
The model covers more than 70 languages, with 25 validated in depth, Japanese among them. It also handles code-switching, where a speaker changes languages mid-sentence. Proper nouns, a known weak spot for speech recognition, can be shored up by passing keywords and context as hints ahead of time.
Best accuracy in third-party testing, but latency is close
Artificial Analysis, an independent evaluator, largely confirmed Meta's claims. English word error rate came in at 3.1 percent, with results settling 0.16 seconds after the speaker stops.
Under the same conditions, ElevenLabs Scribe v2 Realtime scored 3.6 percent at 0.14 seconds, and AssemblyAI Universal-3.5 Pro Realtime scored 4.0 percent. Cartesia Ink-2 splits between 3.4 and 4.0 percent depending on whether it detects utterance endings itself or leans on an external system.
Meta leads on accuracy, but ElevenLabs is marginally faster and the margins are thin. On these numbers, the price gap looks more decisive in practice than the performance edge.
What Meta has not disclosed
Meta has not published the parameter count, the volume of training data, or where the audio came from. The weights are not being released either. Given the company's earlier posture on open publishing, that is hard to read at face value.
Three places can use it today: voice input in Meta AI and Muse Code, and the Meta Model API. Holding the Fn key in any app brings up dictation.
Behind all of this is CEO Mark Zuckerberg's idea of personal superintelligence. In a staged demo, Meta employees argue that dependable speech recognition is the groundwork for personal AI agents that listen in on real conversations through AI glasses. On that front, Germany debated a ban on the company's camera glasses, and the Federal Network Agency decided not to pursue it.
Summary
Muse Voice Transcribe combines transcription, speaker separation, and end-of-speech detection into one model, and balances speed against accuracy by varying its wait time word by word. It takes the top accuracy score in third-party testing, but latency differences are small and performance across vendors is converging. Setting the price at 0.18 USD per hour is where the intent of this launch really shows. The longer the audio you run through it, meeting records and captioning being the obvious cases, the more that gap matters.
