MiniMax H3
MiniMax H3
One-line summary: MiniMax's July 2026 omni-modal video+stereo-audio model; near-tied for first on AA T2V Elo, with partial open weights (base generators public; prompt IR and 2K regen hosted).
What it is
A general-purpose video system that accepts text, images, video, and audio and emits 4–15 s clips with 32 kHz stereo. Official open-source note (2026) splits the stack: H3-Base FL2VA (text / first-last-frame → audio-video) and H3-Base Ref2VA (multi-reference → audio-video) are released; H3-Context-IR and H3-Regenerate-2K stay on the API. From 2026-08-18-autoresearch-best-ai-video-generation-tools.
Why it matters to ai-video-generation
It is the strongest open-weight general video model on the 2026-08-18 AA boards, and Ref2VA is the closest H3 analogue to this thread's reference-driven avatar work. It is not a drop-in wan-animate replacement: motion-mimic papers are not re-run here.
Key facts
- Vendor: MiniMax (Hailuo).
- Architecture: H3-Omni-Transformer, 33B dense; ~13B in AdaLN branches that can be cached at inference.
- Duration: 4–15 s; 24 fps; default shorter side 768 px; 2K via hosted regenerate.
- Dialogue languages (official "stable"): Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish.
- License: MiniMax H3 Community License (not Apache 2.0).
- AA T2V with audio (2026-08-18): Elo 1,237 ±8, $7.80/min — statistically tied with gemini-omni-flash.
- AA I2V with audio: Elo 1,189, second; best open-weights I2V-with-audio.
Strengths (from our perspective)
- Near-frontier Elo at a frontier-low API price.
- Partial self-host path; Ref2VA takes up to 9 images / 3 video clips / 3 audio clips (mixed max 12 files).
Weaknesses (from our perspective)
- "Open" is incomplete: quality-critical IR and 2K regen are hosted.
- Community license, not Apache; US/EU/UK/Korea application form noted on the model card.
- No VRAM floor in the official note — local cost is uncited.