brain/
← all entities
entitygenericartificial-intelligence

Muse Voice Transcribe

Notes

Vintage: 2026-09. Primary evidence is Meta's official post Introducing Muse Voice Transcribe dated September 1, 2026, as hydrated in 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20. Recap treated this as incremental / slightly outside the 24-hour window; the issuer date is inside that window.

Muse Voice Transcribe

One-line summary: Meta Superintelligence Labs' September 1, 2026 real-time audio perception model — streaming ASR, 20+ speaker diarization, endpointing; 70+ languages (25 extensively verified). Recap timing that put it outside the last-24h window is wrong.

What it is

Meta's first real-time audio perception model from Meta Superintelligence Labs. Official post (September 1, 2026): streaming ASR, diarization with 20+ speakers, endpointing; trained with 70+ languages of which 25 are extensively verified; ships via Meta Model API, Meta AI for Mac, and Muse Code.

Why it matters to this thread

Audio / voice models are in-scope under multimodality. The digest recap de-emphasized this as "more incremental or slightly outside the strictest 24-hour window." Issuer date is September 1, 2026 — in-window. File the timing correction; do not adopt the recap's de-emphasis as fact.

Key facts (from 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20)

  • From the same source (Meta official, September 1, 2026): first real-time audio perception model from Meta Superintelligence Labs.
  • From the same source: streaming ASR; diarization with 20+ speakers; endpointing.
  • From the same source: 70+ languages trained, 25 extensively verified.
  • From the same source: ships via Meta Model API, Meta AI for Mac, and Muse Code.

What this source does not establish

Open questions

  • How does Muse Voice Transcribe compare to gemini-3-5-transcribe on noisy IDs / filler-word removal / language coverage? This source does not say.

Sources

Related

Referenced by