The first audio model of Meta's Muse family: a streaming speech-recognition model with speaker diarisation covering more than 70 languages, built on the Muse Spark line and offered through the API. Meta claims a 3.1% streaming word error rate and first place on the Artificial Analysis speech-to-text leaderboard; as of 17 September 2026 that leaderboard shows no Muse row, so the claim is Meta's. No parameter count or pricing is published.

Model Details

License Proprietary (API)
speechaudioproprietary

Related