Hydra-0: Action Flow for Generalist World Modeling and Control

Published
Source
arXiv
Paper number
942
Field
Robotics
arXiv ID
2608.18077

Key points

  • When robot commands are represented as point trajectories in the image plane and supplied as conditions, the world model can predict the effects of movement without knowing each embodiment's coordinate system.
  • The researchers trained one general-purpose world model by expressing heterogeneous videos in a common format, including human first-person demonstrations, hand-gripper footage, and dual-arm and single-arm robot videos.
  • In forward prediction, the method reduced robot-motion error by 90.4% and object-motion error by 60.2%. Its reproduced success rates for five RoboLab policies had a correlation of 0.96.
  • In reverse mode, the model observes only object motion from a human demonstration and generates a compatible robot movement. A trained action head then converts that movement into executable commands.
  • The study evaluated policies only in an open-loop setting and identified closed-loop evaluation as future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)