RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Published
Source
arXiv
Paper number
865
Field
Robotics
arXiv ID
2608.09853

Key points

  • It changes the supervision target for reward models from progress to temporal distance, meaning cost-to-go to the goal, so scaling no longer requires preference labels.
  • It trains on heterogeneous robot data covering 7,000 hours and about 3 million clips.
  • Time-order shuffling and value-isolation attention keep the model from relying only on the learned values.
  • It surpasses the previous best preference-labeled model, 0.655, by reaching 0.675.
  • In real robot RL, it improves success rates from 52.5 percent to 72.5 percent online and from 63.8 percent to 82.5 percent offline.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)