WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 788
- Field
- Robotics
- arXiv ID
- 2607.29613
Key points
- It defines the critic's fundamental limitation in VLA reinforcement learning, when it sees only a single frame, as the state approximation problem, and shows that scalar value regression cannot learn temporal structure.
- WCM is built on a LeJEPA structure and is explicitly trained so that the critic's representations capture temporal dynamics by predicting future latent states and estimating value at the same time.
- It is compatible with modern VLA backbones such as π0, π0.5, and OpenVLA-OFT, and can be integrated into both on-policy and off-policy training.
- It achieves consistently best performance on 149 tasks across four benchmarks, with especially large gains in out-of-distribution generalization.
- On seven real-robot manipulation tasks, it confirms stable performance with OpenVLA-OFT and π0.5-based off-policy RL.
Paper links
External research summaries. These are not HDATF publications or measured product results.