header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Meta's new ASR model achieves real-time transcription, costing only $0.18 per hour.

Perceive Beating AI News Flash: Meta Superintelligence Labs has released Muse Voice Transcribe. It performs real-time speech-to-text conversion, distinguishes between different speakers, and determines when a user has finished speaking, all in a single model. It can process audio files of up to 1 hour in length and differentiate between over 20 individuals.


In an English real-time transcription test by Artificial Analysis, its WER (Word Error Rate, lower is better) was 3.1%. GPT Live Transcribe scored 3.9%, and Gemini 3.5 Transcribe Live scored 4.0%. Muse Voice Transcribe can provide the final text approximately 0.16 seconds after the speaker finishes.


It does not assign the same waiting time to all words. The model reads 80ms of audio each time and decides whether to continue listening or start outputting text. Simple words can be transcribed more quickly, while difficult words are listened to for more context before confirmation. Meta uses reinforcement learning to optimize both accuracy and latency simultaneously.


The model has been trained on over 70 languages, with 25 languages undergoing intensive validation. It can also directly recognize mixed Chinese and English speech in a single sentence. Currently integrated into Meta AI for Mac and Muse Code, it has also opened up the Meta Model API. The pricing is $3 per 1000 minutes, which is equivalent to $0.18 per hour.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish