τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

Published
Source
arXiv
Paper number
736
Field
Robotics
arXiv ID
2607.24485

Key points

  • It learns spatiotemporal tactile representations using a self-supervised objective that predicts future visual features from current touch and subsequent actions.
  • It combines the learned tactile representations with pretrained visual and language features to generate robot actions, while using the prediction branch only during training.
  • It outperformed vision-only VLAs and existing tactile policies on four contact-intensive manipulation tasks, including plug and USB insertion.
  • Removing the additional branch at inference allows it to improve contact-critical robot manipulation without increasing deployment cost.
  • Evaluation focuses on TacAura's four representative tasks and limited task-specific data, leaving validation across a wider variety of robots and contact situations outstanding.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)