Reinforcement Learning for Real-Time Vision-Language-Action Policies

Published
Source
arXiv
Paper number
1089
Field
Robotics
arXiv ID
2609.18207

Key points

  • It solved the problem of stale observations at execution time caused by large VLA model inference latency with a dual structure of 'slow proposal plus fast correction'.
  • Unlike prior real-time execution methods based on imitation learning, it raised reliability beyond the training distribution through reinforcement learning fine-tuning.
  • On four real-robot tasks (object passing, ball balancing, table soccer kicking, dynamic picking), it improved performance from 42% to 97% with only 10 minutes of online data.
  • Training runs fully automatically without human intervention, which matters greatly for real-world deployment.
  • On the Kinetix simulation benchmark it beat all delayed and non-delayed methods, ranking first in all 10 environments.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)