RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

Published
Source
arXiv
Paper number
766
Field
Robotics
arXiv ID
2607.26991

Key points

  • It applies compositional steering with a lightweight RL policy learned on top of VLA latent representations.
  • It finds different scaling laws in success and failure states, and the optimal intervention is to act only on failures.
  • It integrates a SAFE failure detector so that RL steering is activated only when failure is predicted.
  • It achieves up to a 17.3% success-rate improvement on the out-of-distribution SIMPLER and PolaRiS benchmarks.
  • It also transfers to the real PiperX robot, where it improves success by 17.5%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)