Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
- Published
- Source
- arXiv
- Paper number
- 846
- Field
- Computer Vision
- arXiv ID
- 2608.05903
Key points
- It proposed a post-training method that adds semantic information robust to appearance changes while preserving large-scale video pretraining in VAE-based world-action models.
- It inserted training query tokens with temporal positions into the action stream and aligned them with DINOv3 semantic representations of future scenes, keeping additional inference cost low.
- On LIBERO-Plus, it raised out-of-distribution success rates for FastWAM and GE-Act by 9.2 and 8.6 percentage points, respectively, while preserving performance under normal conditions.
- With lighting-color changes on a real Franka robot, it raised GE-Act success from 57.3% to 80.0%, and the same method could be added to multiple world-action architectures.
- The tested changes centered on lighting and color across two simulation benchmarks and one real robot, so robustness to more complex changes in objects, backgrounds, and actions needs further verification.
Paper links
External research summaries. These are not HDATF publications or measured product results.