TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
- Published
- Source
- arXiv
- Paper number
- 328
- Field
- Robotics
- arXiv ID
- 2606.06491
Key points
- Existing VLA models learn only fixed speeds, but TempoVLA controls execution speed with an explicit speed condition.
- VSTA augmentation preserves motion semantics while retiming demonstrations at arbitrary speeds.
- It achieves bidirectional flexible speed control in both simulation and the real world.
- Combined with a large multimodal model, it enables risk-aware dynamic speed control.
- VSTA also improves the base 1x-speed performance by increasing data utilization.
Paper links
External research summaries. These are not HDATF publications or measured product results.