brain/
← all entities
entitygenericartificial-intelligence

rumik-oss 1

Notes

Vintage: 2026-09. Primary evidence is the issuer blog introducing rumik oss 1 and the Hugging Face card, plus non-video team X, as synthesized in 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit. x_video: false — video intro permalink flagged, not a load-bearing cite.

rumik-oss 1

One-line summary: rumik-ai’s 8 Sep 2026 3B open-weights multilingual Indic TTS (4 voices). Issuer claims first open-source Indic TTS with expressive emotions. IndicEmo 2.92/5 vs Gemini 3.1 Flash TTS 4.58 on their table — do not invent “beats all closed models.” License: tiny Aya Fire CC-BY-NC 4.0 + AUP.

What it is

An open-weights Silk-family TTS model. Treat the blog + HF card as fetched issuer grain; treat team X as primary/discourse. Hours figures disagree across team X / blog / HF — keep as an approximate issuer range.

Why it matters to this thread

Audio/voice open-weights and Indic-language coverage are in-scope. Distinct from muse-voice-transcribe (ASR/diarization) and from the same-day fluid-math clip.

Key facts (from 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit)

Issuer blog and HF card

  • From 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit (introducing rumik oss 1, published Sep 8, 2026): first open-source Silk-family model; claims first open-source Indic TTS with expressive emotions; description control (pace/accent/tone) + inline emotion tags; 22 languages; pretrained from tiny Aya Fire + frozen Mimi codec; GRPO post-training.
  • From the same source (same blog): issuer-constructed benches — IndicEmo 2.92/5 overall vs Gemini 3.1 Flash TTS 4.58 on their table; NoVA rendering 0.884.
  • From the same source (huggingface.co/rumik-ai/rumik-oss-1): 3B params; 4 voices (Ira, Aisha, Siya, Zoya); research/non-commercial under tiny Aya Fire CC-BY-NC 4.0 + AUP; long-form >~35s not recommended. HF hours: “fewer than 70,000.” Blog: “66k hours.”

Team X (non-video)

  • From the same source (@lets_dig_deeper): 3B open weights; mid-sentence language switch; research blog + HF links.
  • From the same source (@bhartivatsal): “india’s first open-source expressive multi-indic speech model”; that post says trained on <60k hours (blog/HF say <70k / ~66k — keep as issuer range).

What this source does not establish

Contradictions / tensions

  • Hours range (<60k / 66k / <70k) kept split.
  • Issuer IndicEmo 2.92 vs Gemini 3.1 Flash TTS 4.58 on the same table — Gemini higher; “first open Indic expressive” is an issuer claim, not a closed-model win.

Sources

Related

Referenced by