Pick the video model by the job
As of an August 2026 fetch there is no single winner. Frontier Elo and cheap API pricing can sit at the same end of the table — but iterate, animate a still, long clip, and talking head are different lanes.
Best used to mean expensive. The August 18 fetch of Artificial Analysis’s text-to-video board flipped that. Gemini Omni Flash sat at Elo 1,239 — tied for first — at six dollars for a minute of 1080p. MiniMax H3 was a point behind at 1,237, billed at $7.80. Veo 3.1 Standard, at $0.40 a second, landed mid-pack around 1,089 and $24.00 a minute. Paying more now often buys brand, 4K, or a control surface.
That is a dated snapshot, not a crown. The same table does not pick a talking head, a thirty-second document film, or a clip that still looks like the person you started with. The job does.
What the board actually ranks
Artificial Analysis runs crowdsourced blind-preference Elo for text-to-video, image-to-video, and editing. Votes are pairwise, same prompt, no labels. The price column is standardized: one minute of 1080p at default creator-API settings. That is enough to rank general generation. It is not an avatar bench, a hand bench, or an identity bench.
Vendor blogs still quote last year’s number. Runway’s December 2025 Gen-4.5 post claimed Elo 1,247 and first place. That model is absent from the top twenty-eight text-to-video-with-audio rows fetched in August. MiniMax H3 is the open-weights leader on that board; LTX-2.3 Fast is next. Subscription credit math — HeyGen, Runway Act-Two — is a different unit from the AA minute.
A minute of 1080p, August 18
Artificial Analysis T2V-with-audio, fetched 2026-08-18. Elo is not linear utility; a 150-point gap is not “4× worse.”
Veo Fast, at nine dollars a minute, and Lite, at $4.80, sit in the same Elo band as Standard. OpenAI’s Sora 2 is ten cents a second at 720p; Sora 2 Pro is thirty to seventy cents. Neither had a top-twenty-eight row on the fetched T2V-with-audio table. Official list prices are not the same as dollars per usable clip. OpenAI charges per second requested. Retry rates are uncited.
Iterate, animate, hold the shot
The default generate-and-iterate lane is Omni Flash — official Gemini list about ten cents a second, same model that topped the board. MiniMax H3 is the open-weight near-tie: $7.80 a minute on the AA column, base weights public, the identity-refresh and 2K regeneration hosted. Whether a local H3 matches the hosted Elo without that hosted refresh is open. There is no official VRAM floor.
To animate a still, the image-to-video-with-audio leader on the same fetch was Seedance 2.0, Elo 1,197 at $9.07 a minute. On text-to-video it sat at 1,222, same price. Wan 3.0 is the long-clip / document-to-video pick: official Model Studio prices of five cents a second at 480p, ten at 720p, twenty at 1080p, thirty seconds in a single pass. It has no AA Elo yet. Kling 3.0 list prices did not come back — the site returned a 446.
A talking head is a different machine
Avatar generation splits on the driving signal. Pose-driven models take a reference video, extract a skeleton or a denser body map, and puppet a character through that motion. That is the webcam-influencer lane. Runway Act-Two is the closed product people actually buy; Wan-Animate is the open option that has caught up faster than the audio-driven open lineage. Audio-driven models infer gesture, head movement, and lip sync from a voice track. That is the talking-head lane when you have a podcast and no performance tape. HeyGen Avatar V and OmniHuman sit on that side.
Hybrids are already blurring the line. EchoMimicV2 takes audio first and pose if you have it. Wan 2.2’s speech-to-video model accepts a pose video locked to the audio. HeyGen still needs a reference video for identity; the audio mostly sets rhythm. Calling it “audio-driven” is a little tidy.
Closed audio-driven systems still beat their open counterparts on the May survey. Pose-driven open models closed that gap faster. The board that ranks Omni Flash does not close either comparison. Independent avatar evaluation is asymmetric, not empty.
The benches that are not an avatar board
VBench, from CVPR 2024, is a sixteen-dimension open suite for general text-to-video. One of those dimensions is subject identity inconsistency. VBench++ adds image-to-video. VBench-2.0 adds a Human Fidelity slice. None of the retrieved abstracts names a talking-head task, a pose-driven animation task, or a hand task. Author-reported VBench scores on animation models are uses of a general-video metric. They are not a level-playing-field ranking of avatars.
Talking-head papers have been publishing shared protocols since Chen and colleagues in 2020 — identity, lip-sync, quality, natural motion — because author-run Mechanical Turk was unreproducible. THEval, in 2025, reports 85,000 videos from 17 models and says leaderboards will be released and updated; the 17 names are not in the abstract. Zhang and colleagues, in 2026, put 20 methods across seven datasets under sequence-aligned metrics. Pei and colleagues surveyed talking-face methods on shared datasets as a third-party academic comparison, framed as deepfake research. None of those abstracts names HeyGen Avatar V, Runway Act-Two, or Wan-Animate. “Will be released” is not a live commercial board.
Pose-driven evaluation is thinner. HumanVid states that inaccessible training data hampers fair benchmarking — about 20,000 real 1080p clips plus 10,000 synthetic assets, and an author baseline, not a ranking of MimicMotion or Wan-Animate. HyperMotionX Bench is a public complex-motion set released in the same paper as the authors’ own DiT model. There is no retrieved third-party pose-driven leaderboard. Artificial Analysis remains a general Elo board. The August academic pass did not redo those ranks or prices. The hole that still matters for picking a talking head or a performance take is a live board that includes the products people actually buy.
The face that slowly becomes someone else
Naive video diffusion drifts. The face turns into a cousin. The cloth pattern walks. Proportions shift. Every serious avatar stack has a deliberate identity strategy, and the families do not agree.
One family stuffs a face encoder and an identity adapter into the diffusion loop, then runs an extra face optimization while the image is still denoising — StableAnimator’s claim is that this is end-to-end, no post-hoc face swap. HeyGen skips the embedding bottleneck and conditions on the full token sequence of the reference video at every transformer layer, with sparsity so cost grows near-linearly with reference length. It tries to keep dental structure and skin texture in one bucket and talking rhythm in another. A third family bakes the person into a 3D body model and Gaussian splat — diffusion is only a synthetic-data factory — then oversamples the input image so the bake does not wander. Wan-Animate encodes raw face images, not landmarks, into dedicated face blocks, and uses a spatially aligned skeleton for the body.
Full-token attention is the expensive, high-fidelity story, especially for micro-expression and idiosyncratic talking style. The 3D bake trades about an hour of up-front compute per subject for cheap frames afterward. Those are architectural claims from the May 2026 survey. HeyGen’s pairwise-preference percentages and Wan-Animate’s author-reported wins over Act-Two are vendor numbers. They are not an independent ranking.
Under the hood: UNet gave way to DiT
Through 2025 the avatar stack migrated from UNet diffusion — Stable Video Diffusion, EchoMimicV2’s denoising_unet checkpoints, first-generation AnimateAnyone — to Diffusion Transformer backbones. Wan-Animate on Wan-I2V, UniAnimate-DiT on Wan 2.1-14B, and the Wan 2.2 family sit on the DiT side. Wan-Animate’s authors claim the transformer backbone buys superior temporal coherence versus UNet-based AnimateAnyone V2; that is their paper, not an independent bench.
For new work in 2026, defaulting to a DiT backbone — usually Wan — is the lower-risk architectural bet. UNet models still ship and work; the coherence ceiling is lower and compute tradeoffs differ. LoRA fine-tuning on Wan I2V, as UniAnimate-DiT does, is one cheap path onto the new spine. Whether audio-driven lines like EchoMimicV2 follow the same migration is still open; pose-driven avatars moved first.