SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction

Published
Source
arXiv
Paper number
517
Field
Robotics
arXiv ID
2606.27581

Key points

  • Using reference motion plus per-link contact labels as policy conditioning, we unify free-space and contact-rich behaviors in a single policy for the first time.
  • Hindsight scene reconstruction infers the scene-interaction graph from retargeted human motion and reconstructs terrain and object assets in reverse.
  • With 7.5 hours of contact-rich training data, combining AMASS, OMOMO, Bones, and Lafan, it generalizes to unseen motions and environments.
  • Global root tracking is essential: without it, terrain success drops from 100% to 45% and object-task success from 95% to 20%.
  • Removing contact labels makes object-grasp success converge to 0%, proving the decisive role of contact labels.
  • On the real Unitree G1, state-estimation performance is confirmed with an average root-position error of 0.032 m and a rotation error of 0.018 rad.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)