BadWAM: When World-Action Models Dream Right but Act Wrong
- Published
- Source
- arXiv
- Paper number
- 654
- Field
- Machine Learning
- arXiv ID
- 2607.15207
Key points
- We are the first to uncover a new vulnerability in WAMs, showing that imagination and action can be separated.
- An attack that targets action only reduces robotic task success from 96.5% to 43.1%, cutting it by more than half.
- We design a stealth attack that preserves imagination while quietly corrupting only action, demonstrating that existing safety monitors can be bypassed.
- The attack transfers from one WAM variant to another, confirming that the vulnerability is not specific to a single implementation.
- Simple preprocessing defenses are not enough, which suggests a new safety approach that directly checks synchronization between action and imagination.
Paper links
External research summaries. These are not HDATF publications or measured product results.