RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
- Published
- Source
- arXiv
- Paper number
- 865
- Field
- Robotics
- arXiv ID
- 2608.09853
Key points
- It changes the supervision target for reward models from progress to temporal distance, meaning cost-to-go to the goal, so scaling no longer requires preference labels.
- It trains on heterogeneous robot data covering 7,000 hours and about 3 million clips.
- Time-order shuffling and value-isolation attention keep the model from relying only on the learned values.
- It surpasses the previous best preference-labeled model, 0.655, by reaching 0.675.
- In real robot RL, it improves success rates from 52.5 percent to 72.5 percent online and from 63.8 percent to 82.5 percent offline.
Paper links
External research summaries. These are not HDATF publications or measured product results.