World Value Models for Robotic Manipulation
- Published
- Source
- arXiv
- Paper number
- 488
- Field
- Robotics
- arXiv ID
- 2606.24742
Key points
- This is the first approach to reuse the spatiotemporal prior knowledge of world models as the basis for robotic value estimation.
- It combines a video-generation stream and a value DiT through a Mixture-of-Transformers, or MoT, to minimize representation interference.
- It formulates the value function as a flow-matching-trained distributional chunk, which provides dense learning signals.
- It releases Suboptimal-Value-Bench, which consists of 800 suboptimal trajectories plus human labels.
- It achieves SOTA on both the expert VOC and Suboptimal-Value-Bench, and it also improves downstream policy learning consistently.
Paper links
External research summaries. These are not HDATF publications or measured product results.