RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
- Published
- Source
- arXiv
- Paper number
- 766
- Field
- Robotics
- arXiv ID
- 2607.26991
Key points
- It applies compositional steering with a lightweight RL policy learned on top of VLA latent representations.
- It finds different scaling laws in success and failure states, and the optimal intervention is to act only on failures.
- It integrates a SAFE failure detector so that RL steering is activated only when failure is predicted.
- It achieves up to a 17.3% success-rate improvement on the out-of-distribution SIMPLER and PolaRiS benchmarks.
- It also transfers to the real PiperX robot, where it improves success by 17.5%.
Paper links
External research summaries. These are not HDATF publications or measured product results.