BadWAM: When World-Action Models Dream Right but Act Wrong

Published
Source
arXiv
Paper number
654
Field
Machine Learning
arXiv ID
2607.15207

Key points

  • We are the first to uncover a new vulnerability in WAMs, showing that imagination and action can be separated.
  • An attack that targets action only reduces robotic task success from 96.5% to 43.1%, cutting it by more than half.
  • We design a stealth attack that preserves imagination while quietly corrupting only action, demonstrating that existing safety monitors can be bypassed.
  • The attack transfers from one WAM variant to another, confirming that the vulnerability is not specific to a single implementation.
  • Simple preprocessing defenses are not enough, which suggests a new safety approach that directly checks synchronization between action and imagination.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)