ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Published
Source
arXiv
Paper number
455
Field
Computer Vision
arXiv ID
2606.19531

Key points

  • It points out three limitations of video-generation WAMs, namely cost, irrelevant details, and error accumulation, and proposes image editing as an alternative.
  • It passes the KV cache from the editing backbone to a flow-matching action expert, enabling action prediction without image decoding.
  • It achieves a 98.4% success rate on LIBERO and an 84.5% success rate on real robots, outperforming π0 at 55.8% and FastWAM at 79.0%.
  • It achieves FLOPs reduction from 63.65 to 9.72, a 6x reduction, and latency reduction from 1081 ms to 263 ms, about 4x faster.
  • It can use a variety of editing backbones such as FLUX.2, OmniGen2, and Ovis-U1, so it is not tied to a specific model.
  • Attention analysis confirms that the editing cache focuses on task-relevant changed regions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)