Talking-head evaluation benchmarks
Talking-head evaluation benchmarks
One-line summary: Academic talking-head generation has dedicated, shared-protocol benches (identity, lip-sync, quality, motion) — including a 2025 framework that promises a public leaderboard — but retrieved abstracts do not name HeyGen, Act-Two, or Wan-Animate on those boards.
The insight
The audio-driven / talking-face lane is not evaluation-empty. The gap the thread asked about is sharper than "no independent benches exist": talking-head papers have been publishing standardized protocols since 2020, and 2025–2026 work claims multi-model, updatable leaderboards. What the abstracts do not show is a live commercial-inclusive board for the closed products this thread actually compares.
Evidence
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: Chen et al. (ArXiv 2020) publish a talking-head benchmark with standardized pre-processing and metrics for identity preserving, lip synchronization, video quality, and natural-spontaneous motion, and they criticize unreproducible author-run MTurk evals.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: THEval (ArXiv 2025) proposes 8 metrics on quality, naturalness, and synchronization; reports 85,000 videos from 17 models; says code, dataset, and leaderboards will be publicly released and regularly updated. The 17 models are not named in the retrieved abstract.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: Zhang et al. (ArXiv 2026) recast talking-head eval as sequence alignment (Soft-DTW) and report 20 methods across seven datasets under standardized protocols. Methods not named; 0 citations in that retrieval.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: Pei et al. (ACM Computing Surveys 2024) benchmark representative talking-face / reenactment methods on widely adopted datasets as a third-party academic comparison. Framed as deepfake / digital-human research, not a commercial avatar product board.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: Yaman et al. (ArXiv 2025) give a model-agnostic identity-leakage (lip leakage) protocol; methodology paper, not a leaderboard.
Design implications
When comparing talking-head academic methods, prefer Chen / THEval / Zhang-style shared protocols over author MTurk. When picking heygen-avatar-v vs runway-act-two, this literature has not yet been shown (in retrieved abstracts) to include those names. See independent-avatar-benchmarks.
Contradictions / tensions
- THEval and Zhang claim updatable or large-scale multi-method boards; both are low-citation ArXiv in the 2026-08-21 retrieval. "Will be released" is not the same as a live board.
- Identity on these protocols is a talking-face / speaker property. It is not shown to be the same measurement as vbench subject identity inconsistency.
Open questions
- Which 17 / 20 methods sit on THEval and Zhang's boards, and are any commercial?
- Is the THEval leaderboard actually public yet? Abstracts only; no landing page was fetched.