TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies

Published
Source
arXiv
Paper number
328
Field
Robotics
arXiv ID
2606.06491

Key points

  • Existing VLA models learn only fixed speeds, but TempoVLA controls execution speed with an explicit speed condition.
  • VSTA augmentation preserves motion semantics while retiming demonstrations at arbitrary speeds.
  • It achieves bidirectional flexible speed control in both simulation and the real world.
  • Combined with a large multimodal model, it enables risk-aware dynamic speed control.
  • VSTA also improves the base 1x-speed performance by increasing data utilization.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)