HumanVid
HumanVid
One-line summary: 2024 camera-controllable human-image-animation dataset whose abstract states that inaccessible training data hampers fair and transparent benchmarking.
What it is
A large-scale human-animation dataset: about 20K filtered 1080p real-world human-centric videos plus about 10K synthetic 3D-avatar assets with rule-based camera trajectories. Introduces a CamAnimate baseline that conditions on both human pose and camera motion. From 2026-08-21-academic-research-independent-avatar-generation-benchmarks.
Why it matters to ai-video-generation
This is the clearest retrieved academic statement of the pose-driven fairness gap this thread already suspected: impressive author results rest on private high-quality training data, so head-to-head numbers are not on a level playing field. HumanVid is a public dataset plus an author baseline — not a third-party ranking of mimicmotion, stableanimator, or wan-animate.
Key facts
- Real split: 20K 1080p videos; 2D pose + SLAM-based camera annotation.
- Synthetic split: 10K 3D avatar assets; rule-based camera trajectories.
- Baseline: CamAnimate (camera-controllable human animation).
- Project URL named in the abstract: https://humanvid.github.io/ (not fetched this pass).
Strengths
- Names the private-data benchmarking problem in print.
- Public data + code claim.
Weaknesses
- Author baseline "sets a new benchmark" is still an author-reported result.
- Does not, in the retrieved abstract, rank prior pose-driven SOTA on a shared protocol.