Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

Published
Source
arXiv
Paper number
1033
Field
Robotics
arXiv ID
2608.23831

Key points

  • It identified the cause of failure: VLA inference latency of 100–300ms changes environment dynamics and breaks RL's Markov assumption.
  • It restored an approximately Markov structure by augmenting the state with committed actions and observations collected during inference.
  • It enabled stable fine-tuning even under latency conditions where standard RL failed completely.
  • It matched or exceeded standard RL's performance under ideal conditions with no latency.
  • It demonstrated real-world improvements under asynchronous inference on 3 tasks using a real dual-arm UR5e robot.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)