Independent evaluation by Artificial Analysis confirms Meta's claims: Muse Voice Transcribe achieves a 3.1 percent word error rate on English in 0.16 seconds. The Spark-family model, Meta Superintelligence Labs' first real-time audio perception model, distinguishes over 20 speakers and supports more than 70 languages. At $0.18 per hour, it undercuts OpenAI and ElevenLabs pricing. It is available now in Meta AI and via the Meta Model API; Meta has not disclosed parameter counts or released weights.
No score is assigned. Sources and their independence are shown in the citation chain below.