World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

Published
Source
arXiv
Paper number
519
Field
Robotics
arXiv ID
2606.27374

Key points

  • The first continual imitation learning approach to use WAM generative capability, reconstructed future frames, as pseudo-replay.
  • Compared with sequential fine-tuning, it reduces catastrophic forgetting by up to 50% and comes close to experience replay that uses real demonstrations.
  • Recursive generation feeds future observations predicted by the WAM back into the model to synthesize complete trajectories, which is recurrent generative replay.
  • It remains effective on real robots as well: NBT improves from 96.3 to 60.5, reducing forgetting by 40%, and FWT improves from 50 to 80.
  • It identifies two main bottlenecks: visual quality degradation during long-horizon recursive generation, reflected in lower PSNR, and mismatch between imagined and grounded actions, 83% versus 42%.
  • Action representation drift shows that REGEN, at 0.12, preserves representations much better than Seq-FT, at 0.30.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)