$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Published
Source
arXiv
Paper number
834
Field
Robotics
arXiv ID
2608.06375

Key points

  • We trained full-body cooperative motion in one model without separating locomotion and manipulation.
  • Instead of future video generation, we improved action quality using latent embedding prediction, a lightweight world model.
  • We built and released omega-HOME, a 40-hour real-home humanoid dataset.
  • It performs 11 household tasks with a single model and outperforms prior imitation learning, VLA, and WAM methods.
  • We propose a pipeline that converts human video data into robot-executable actions through simulation replay.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)