GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

Published
Source
arXiv
Paper number
308
Field
Robotics
arXiv ID
2606.05160

Key points

  • This privileged setting better conditions 4D reconstruction, allowing model-based object tracking, human motion estimation, and interaction-recognition optimization to reconstruct metric 4D human-object interaction (HOI) trajectories while reducing depth ambiguity and shape mismatch.
  • The authors retarget the reconstructed motions to a humanoid robot and learn complementary general-purpose task trackers: an object-aware latent adapter for manipulation and a scene-aware tracker for terrain traversal.
  • Using only data generated by GRAIL, they learn egocentric visual policies through a sim-to-real pipeline and deploy them on the Unitree G1 humanoid, achieving real-world success rates of 84% on diverse object picking tasks and 90% on stair climbing.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)