Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency
- Published
- Source
- arXiv
- Paper number
- 1033
- Field
- Robotics
- arXiv ID
- 2608.23831
Key points
- It identified the cause of failure: VLA inference latency of 100–300ms changes environment dynamics and breaks RL's Markov assumption.
- It restored an approximately Markov structure by augmenting the state with committed actions and observations collected during inference.
- It enabled stable fine-tuning even under latency conditions where standard RL failed completely.
- It matched or exceeded standard RL's performance under ideal conditions with no latency.
- It demonstrated real-world improvements under asynchronous inference on 3 tasks using a real dual-arm UR5e robot.
Paper links
External research summaries. These are not HDATF publications or measured product results.