World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays
- Published
- Source
- arXiv
- Paper number
- 519
- Field
- Robotics
- arXiv ID
- 2606.27374
Key points
- The first continual imitation learning approach to use WAM generative capability, reconstructed future frames, as pseudo-replay.
- Compared with sequential fine-tuning, it reduces catastrophic forgetting by up to 50% and comes close to experience replay that uses real demonstrations.
- Recursive generation feeds future observations predicted by the WAM back into the model to synthesize complete trajectories, which is recurrent generative replay.
- It remains effective on real robots as well: NBT improves from 96.3 to 60.5, reducing forgetting by 40%, and FWT improves from 50 to 80.
- It identifies two main bottlenecks: visual quality degradation during long-horizon recursive generation, reflected in lower PSNR, and mismatch between imagined and grounded actions, 83% versus 42%.
- Action representation drift shows that REGEN, at 0.12, preserves representations much better than Seq-FT, at 0.30.
Paper links
External research summaries. These are not HDATF publications or measured product results.