τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
- Published
- Source
- arXiv
- Paper number
- 736
- Field
- Robotics
- arXiv ID
- 2607.24485
Key points
- It learns spatiotemporal tactile representations using a self-supervised objective that predicts future visual features from current touch and subsequent actions.
- It combines the learned tactile representations with pretrained visual and language features to generate robot actions, while using the prediction branch only during training.
- It outperformed vision-only VLAs and existing tactile policies on four contact-intensive manipulation tasks, including plug and USB insertion.
- Removing the additional branch at inference allows it to improve contact-critical robot manipulation without increasing deployment cost.
- Evaluation focuses on TacAura's four representative tasks and limited task-specific data, leaving validation across a wider variety of robots and contact situations outstanding.
Paper links
External research summaries. These are not HDATF publications or measured product results.