FACT: Failure-Aware Causal Training for World-Action Models

Published
Source
arXiv
Paper number
892
Field
Robotics
arXiv ID
2608.10232

Key points

  • It predicts future video and progress conditioned on actions created first, and it uses failed actions only for outcome prediction, not imitation.
  • On 50 RoboTwin tasks, average success is 81.8% without video co-training, 85.6% with video co-training, and 87.5% when failure data are added.
  • On a real dual-arm robot, tasks seen during training improve from 82% to 89%, and unseen variants rise from 67% to 77%.
  • Video prediction PSNR on failure cases improves from 19.51 to 25.92, while success cases stay almost unchanged at 26.12 and 26.08, so normal prediction is preserved.
  • Evaluating four action candidates by progress gives 92% on seen tasks and 82% on new variants, but candidate evaluation requires extra computation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)