ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

Published
Source
arXiv
Paper number
724
Field
Robotics
arXiv ID
2607.22530

Key points

  • It proposes the first action-conditioned visual and tactile world model for robots and builds a training pipeline that unifies simulation and real-world data.
  • By modeling touch as an additional view, it uses view-aware conditioning and cross-view attention to preserve visual-tactile consistency.
  • Adding generated rollouts to policy training improves the average success rate of the pi0.5 plus tactile policy from 42.5 percent to 67.5 percent, a gain of 25 points.
  • Pretraining and simulation data greatly improve generation quality across PSNR, SSIM, and LPIPS.
  • It can also be used as a policy evaluation tool to predict visual and tactile outcomes for a given action sequence.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)