conceptai-video-generation
Artificial Analysis Video Arena
Notes
Artificial Analysis Video Arena
One-line summary: Crowdsourced blind-preference Elo boards for text-to-video, image-to-video, and video editing, with a creator-API $/min column — the independent ranking used in the 2026-08-18 tools pass.
The insight
Vendor blogs still quote last year's Elo. The live AA tables are a different object: pairwise votes on the same prompt, no model labels, plus a standardized "1 minute of 1080p at default settings" price. That is enough to rank general video tools. It is not an avatar, hand, or identity bench.
Evidence
- From 2026-08-18-autoresearch-best-ai-video-generation-tools: T2V-with-audio leaders on 2026-08-18 were gemini-omni-flash (Elo 1,239, $6.00/min) and minimax-h3 (1,237, $7.80/min); I2V-with-audio leader was seedance-2-0 (1,197, $9.07/min).
- From 2026-08-18-autoresearch-best-ai-video-generation-tools: AA's FAQ states ranks come from blind user votes; open-weights T2V-with-audio leader is MiniMax H3, then LTX-2.3 Fast.
- From 2026-08-18-autoresearch-best-ai-video-generation-tools: Runway's 1 Dec 2025 Gen-4.5 post claimed Elo 1,247 / #1 on AA T2V; that model is absent from the top-28 T2V-with-audio table fetched 2026-08-18.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: the academic VBench-class suites (vbench) are a second general-video evaluation object, also not an avatar/identity/hand board. That pass did not redo AA Elo or prices.
Design implications
Use AA to pick a default T2V/I2V API. Do not use it to close independent-avatar-benchmarks or hand-fidelity-comparison-across-avatar-models.
Contradictions / tensions
- Stale vendor Elo (Runway Gen-4.5 1,247) vs live table.
- Price column is creator-API default 1080p — subscription credit math (runway-act-two, heygen-avatar-v) is a different unit.
Open questions
- Does AA publish a historical series that would explain Gen-4.5's disappearance?
- Any avatar-specific AA slice?
Related
Referenced by