OMG: Omni-Modal Motion Generation for Generalist Humanoid Control

Published
Source
arXiv
Paper number
384
Field
Robotics
arXiv ID
2606.10340

Key points

  • A hierarchical structure with a motion-generation brain and a motion-tracking cerebellum generalizes whole-body humanoid control.
  • OMG-Data is a 1,174.66-hour omni-modal humanoid motion corpus containing language, audio, and human motion.
  • OMG-DiT is a diffusion transformer conditioned on language, audio, and human references, and it can be extended to new modalities.
  • It achieves an FID of 6.03 on text-to-motion with OMG-XL, which is a large improvement over prior methods.
  • It supports zero-shot compositionality, so unseen combinations of language and audio can be combined at inference time.
  • Performance improves consistently as the model scales from B to L to XL.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)