ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
- Published
- Source
- arXiv
- Paper number
- 455
- Field
- Computer Vision
- arXiv ID
- 2606.19531
Key points
- It points out three limitations of video-generation WAMs, namely cost, irrelevant details, and error accumulation, and proposes image editing as an alternative.
- It passes the KV cache from the editing backbone to a flow-matching action expert, enabling action prediction without image decoding.
- It achieves a 98.4% success rate on LIBERO and an 84.5% success rate on real robots, outperforming π0 at 55.8% and FastWAM at 79.0%.
- It achieves FLOPs reduction from 63.65 to 9.72, a 6x reduction, and latency reduction from 1081 ms to 263 ms, about 4x faster.
- It can use a variety of editing backbones such as FLUX.2, OmniGen2, and Ovis-U1, so it is not tied to a specific model.
- Attention analysis confirms that the editing cache focuses on task-relevant changed regions.
Paper links
External research summaries. These are not HDATF publications or measured product results.