ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
- Published
- Source
- arXiv
- Paper number
- 787
- Field
- Robotics
- arXiv ID
- 2607.28993
Key points
- The authors discover a hallucination phenomenon in which world action models hallucinate training-scene content in visually unfamiliar environments, and they confirm that 70.6 percent of 180 inspected cases exhibit this behavior.
- They show through 290-frame triplet experiments that DINOv3 features are more robust to visual changes than VAE latent representations while still separating task states well.
- A dual-space future expert predicts both VAE and DINO spaces, and a current-anchor intention retrieval module finds task-relevant evidence from the recent DINO history.
- It achieves 98.7 percent on LIBERO and 92.8 percent on RoboTwin 2.0, and on visually perturbed environments such as LIBERO-Plus it improves by 21.3 points over Fast-WAM and raises real-robot success from 25.8 percent to 61.5 percent.
- It is trained end to end without extra pretraining or task-specific annotation, and it is efficient because it does not require explicit future generation at inference time.
Paper links
External research summaries. These are not HDATF publications or measured product results.