brain/
← all entities
entitygenericai-video-generation

HumanVid

Notes

HumanVid

One-line summary: 2024 camera-controllable human-image-animation dataset whose abstract states that inaccessible training data hampers fair and transparent benchmarking.

What it is

A large-scale human-animation dataset: about 20K filtered 1080p real-world human-centric videos plus about 10K synthetic 3D-avatar assets with rule-based camera trajectories. Introduces a CamAnimate baseline that conditions on both human pose and camera motion. From 2026-08-21-academic-research-independent-avatar-generation-benchmarks.

Why it matters to ai-video-generation

This is the clearest retrieved academic statement of the pose-driven fairness gap this thread already suspected: impressive author results rest on private high-quality training data, so head-to-head numbers are not on a level playing field. HumanVid is a public dataset plus an author baseline — not a third-party ranking of mimicmotion, stableanimator, or wan-animate.

Key facts

  • Real split: 20K 1080p videos; 2D pose + SLAM-based camera annotation.
  • Synthetic split: 10K 3D avatar assets; rule-based camera trajectories.
  • Baseline: CamAnimate (camera-controllable human animation).
  • Project URL named in the abstract: https://humanvid.github.io/ (not fetched this pass).

Strengths

  • Names the private-data benchmarking problem in print.
  • Public data + code claim.

Weaknesses

  • Author baseline "sets a new benchmark" is still an author-reported result.
  • Does not, in the retrieved abstract, rank prior pose-driven SOTA on a shared protocol.

Sources

Related

Referenced by