RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
- Published
- Source
- arXiv
- Paper number
- 416
- Field
- Computer Vision
- arXiv ID
- 2606.13674
Key points
- We identify a problem with the reconstruction-oriented tokenizer in existing WAM: it misses manipulation semantics, so we propose a semantic visual-action tokenizer.
- RepViTok aligns video autoencoder latents with DINOv2 and induces latent actions in the same space.
- A two-stage training process, pretraining on visual plus latent actions and then adapting to robot behaviors, outperforms joint prediction.
- On RoboTwin 2.0, it reaches 86.6 on Easy and 83.1 on Hard, a large improvement over the WAN2.2 VAE-based baseline across an average of 50 tasks.
- Its best performance occurs at CFG scale 1.0, showing that semantic alignment greatly reduces dependence on video CFG.
- In latent-action visualizations, it shows patterns more focused on manipulation-related changes than LAPA does.
Paper links
External research summaries. These are not HDATF publications or measured product results.