World Value Models for Robotic Manipulation

Published
Source
arXiv
Paper number
488
Field
Robotics
arXiv ID
2606.24742

Key points

  • This is the first approach to reuse the spatiotemporal prior knowledge of world models as the basis for robotic value estimation.
  • It combines a video-generation stream and a value DiT through a Mixture-of-Transformers, or MoT, to minimize representation interference.
  • It formulates the value function as a flow-matching-trained distributional chunk, which provides dense learning signals.
  • It releases Suboptimal-Value-Bench, which consists of 800 suboptimal trajectories plus human labels.
  • It achieves SOTA on both the expert VOC and Suboptimal-Value-Bench, and it also improves downstream policy learning consistently.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)