VBench
VBench
One-line summary: Open, multi-dimension benchmark family for general text-to-video and image-to-video quality — independent of any one model author, but not an avatar / talking-head / hand task board in the retrieved abstracts.
The insight
VBench-class suites exist and are independent. They dissect video quality into named dimensions (including subject identity inconsistency) and, later, intrinsic-faithfulness slices such as Human Fidelity. That is not the same object as a level-playing-field ranking of pose-driven or talking-head avatar models. Author-reported VBench scores on animation models are uses of a general-video metric.
Evidence
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: VBench (CVPR 2024) has 16 dimensions, among them subject identity inconsistency, motion smoothness, temporal flickering, and spatial relationship; human-preference annotations per dimension; prompts, methods, videos, and annotations open-sourced.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: VBench++ (TPAMI 2025) keeps the 16 T2V dimensions and adds an I2V Image Suite plus trustworthiness evaluation; fully open-sourced.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: an ArXiv write-up of VBench++ says the authors continually add video generation models to a leaderboard — described as T2V/I2V, not avatar.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: VBench-2.0 (ArXiv 2025) adds five intrinsic-faithfulness dimensions including Human Fidelity (anatomical correctness among the motives); the abstract does not name talking-head, pose-driven animation, or hand tasks.
- From 2026-08-21-academic-research-independent-avatar-generation-benchmarks: WorldJen (ArXiv 2026) is a later general-video bench that still evaluates T2V models, not avatars. This page does not restated its model ranks.
Design implications
Use VBench to interpret general video-quality claims. Do not treat a VBench number, or VBench-2.0 Human Fidelity, as an independent avatar / identity / hand leaderboard. See independent-avatar-benchmarks.
Contradictions / tensions
- dispose-pose-conditioning reports VBench improvements over MimicMotion from the May 2026 survey. That is author-applied general-video scoring, not a VBench avatar task suite. From 2026-05-07-ai-avatar-motion-mimicking-models-survey and 2026-08-21-academic-research-independent-avatar-generation-benchmarks.
Open questions
- Does any VBench release (beyond the retrieved abstracts) define a human-animation or talking-head task?
- Is VBench-2.0 Human Fidelity a face/body identity metric or T2V anatomical correctness?