SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

Published
Source
arXiv
Paper number
450
Field
Robotics
arXiv ID
2606.18610

Key points

  • Joint training of forward and backward dynamics corrects autoregressive rollout drift using backward signals.
  • Cross-view inpainting training preserves multi-camera consistency without explicit memory.
  • The backward dynamics mode is reused as a per-chunk uncertainty signal for early stopping when drift appears.
  • Across seven real VLA policies, it achieves a closed-loop Pearson correlation of 0.929 and MMRV of 0.119, outperforming Ctrl-World, IRASim, and Cosmos-Predict 2.5.
  • It classifies the real failure modes of policies, including language understanding, grasping, and releasing, into four categories and achieves more than 70 percent recall.
  • It is trained on 381 hours of real table-bussing data and shows generalization to out-of-distribution tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)