RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

Published
Source
arXiv
Paper number
416
Field
Computer Vision
arXiv ID
2606.13674

Key points

  • We identify a problem with the reconstruction-oriented tokenizer in existing WAM: it misses manipulation semantics, so we propose a semantic visual-action tokenizer.
  • RepViTok aligns video autoencoder latents with DINOv2 and induces latent actions in the same space.
  • A two-stage training process, pretraining on visual plus latent actions and then adapting to robot behaviors, outperforms joint prediction.
  • On RoboTwin 2.0, it reaches 86.6 on Easy and 83.1 on Hard, a large improvement over the WAN2.2 VAE-based baseline across an average of 50 tasks.
  • Its best performance occurs at CFG scale 1.0, showing that semantic alignment greatly reduces dependence on video CFG.
  • In latent-action visualizations, it shows patterns more focused on manipulation-related changes than LAPA does.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)