OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Published
Source
arXiv
Paper number
026
Field
Video Generation
arXiv ID
2502.01061

Key points

  • End-to-end human animation models are hard to scale to large datasets because strict data filtering requirements discard more than 90% of the raw video data.
  • Weak correlation between major driving signals, such as audio for lip sync, and general body motion restricts existing models to specialized scenarios and blocks full-body animation and diverse interactions.
  • Current models are limited to narrow use cases such as front-facing portraits with static backgrounds because small, highly curated datasets limit generalization.
  • OmniHuman is proposed as an end-to-end multimodal conditional human video generation framework built on a more advanced Diffusion Transformer (DiT) backbone.
  • It adopts a new Omni-conditions training strategy with a progressive three-stage mixed-conditioning scheme that expands data with weak conditions while carefully balancing strong versus weak conditioning during training.
  • It efficiently integrates diverse driving conditions, such as audio, pose, and text, and appearance conditions, such as reference images and motion frames, with minimal additional parameters, and it reuses the DiT backbone for appearance conditioning.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)