WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Published
Source
arXiv
Paper number
788
Field
Robotics
arXiv ID
2607.29613

Key points

  • It defines the critic's fundamental limitation in VLA reinforcement learning, when it sees only a single frame, as the state approximation problem, and shows that scalar value regression cannot learn temporal structure.
  • WCM is built on a LeJEPA structure and is explicitly trained so that the critic's representations capture temporal dynamics by predicting future latent states and estimating value at the same time.
  • It is compatible with modern VLA backbones such as π0, π0.5, and OpenVLA-OFT, and can be integrated into both on-policy and off-policy training.
  • It achieves consistently best performance on 149 tasks across four benchmarks, with especially large gains in out-of-distribution generalization.
  • On seven real-robot manipulation tasks, it confirms stable performance with OpenVLA-OFT and π0.5-based off-policy RL.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)