ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation
- Published
- Source
- arXiv
- Paper number
- 724
- Field
- Robotics
- arXiv ID
- 2607.22530
Key points
- It proposes the first action-conditioned visual and tactile world model for robots and builds a training pipeline that unifies simulation and real-world data.
- By modeling touch as an additional view, it uses view-aware conditioning and cross-view attention to preserve visual-tactile consistency.
- Adding generated rollouts to policy training improves the average success rate of the pi0.5 plus tactile policy from 42.5 percent to 67.5 percent, a gain of 25 points.
- Pretraining and simulation data greatly improve generation quality across PSNR, SSIM, and LPIPS.
- It can also be used as a policy evaluation tool to predict visual and tactile outcomes for a given action sequence.
Paper links
External research summaries. These are not HDATF publications or measured product results.