ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

Published
Source
arXiv
Paper number
787
Field
Robotics
arXiv ID
2607.28993

Key points

  • The authors discover a hallucination phenomenon in which world action models hallucinate training-scene content in visually unfamiliar environments, and they confirm that 70.6 percent of 180 inspected cases exhibit this behavior.
  • They show through 290-frame triplet experiments that DINOv3 features are more robust to visual changes than VAE latent representations while still separating task states well.
  • A dual-space future expert predicts both VAE and DINO spaces, and a current-anchor intention retrieval module finds task-relevant evidence from the recent DINO history.
  • It achieves 98.7 percent on LIBERO and 92.8 percent on RoboTwin 2.0, and on visually perturbed environments such as LIBERO-Plus it improves by 21.3 points over Fast-WAM and raises real-robot success from 25.8 percent to 61.5 percent.
  • It is trained end to end without extra pretraining or task-specific annotation, and it is efficient because it does not require explicit future generation at inference time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)