Reinforcement Learning for Real-Time Vision-Language-Action Policies
- Published
- Source
- arXiv
- Paper number
- 1089
- Field
- Robotics
- arXiv ID
- 2609.18207
Key points
- It solved the problem of stale observations at execution time caused by large VLA model inference latency with a dual structure of 'slow proposal plus fast correction'.
- Unlike prior real-time execution methods based on imitation learning, it raised reliability beyond the training distribution through reinforcement learning fine-tuning.
- On four real-robot tasks (object passing, ball balancing, table soccer kicking, dynamic picking), it improved performance from 42% to 97% with only 10 minutes of online data.
- Training runs fully automatically without human intervention, which matters greatly for real-world deployment.
- On the Kinetix simulation benchmark it beat all delayed and non-delayed methods, ranking first in all 10 environments.
Paper links
External research summaries. These are not HDATF publications or measured product results.