SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
- Published
- Source
- arXiv
- Paper number
- 517
- Field
- Robotics
- arXiv ID
- 2606.27581
Key points
- Using reference motion plus per-link contact labels as policy conditioning, we unify free-space and contact-rich behaviors in a single policy for the first time.
- Hindsight scene reconstruction infers the scene-interaction graph from retargeted human motion and reconstructs terrain and object assets in reverse.
- With 7.5 hours of contact-rich training data, combining AMASS, OMOMO, Bones, and Lafan, it generalizes to unseen motions and environments.
- Global root tracking is essential: without it, terrain success drops from 100% to 45% and object-task success from 95% to 20%.
- Removing contact labels makes object-grasp success converge to 0%, proving the decisive role of contact labels.
- On the real Unitree G1, state-estimation performance is confirmed with an average root-position error of 0.032 m and a rotation error of 0.018 rad.
Paper links
External research summaries. These are not HDATF publications or measured product results.