🛰️ Daily AI Frontier
‹ back to 2026-07-16

BadWAM: When World-Action Models Dream Right but Act Wrong

arXiv cs.LG Multimodal & Generative Qi Li, Xingyi Yang, Xinchao Wang 2026-07-16

TL;DR - BadWAM introduces a class of adversarial attacks against world-action models (WAMs) for embodied robot control, showing that the assumed safety benefit of coupling action generation with future-world prediction is fragile: tiny visual perturbations can desynchronize what a robot imagines from what it actually does.

  • Defines "World-Action Drift Attacks" along two axes—attack strength and stealthiness—unifying overt and covert failure modes.
  • An action-only attack maximizes disruption, cutting task success from 96.5% to 43.1% under closed-loop execution.
  • An imagination-preserving attack stays stealthy: it induces harmful action shifts while keeping the predicted future close to the clean imagination, so the model "dreams right but acts wrong."
  • Shows moderate future-preserving regularization retains strong attack performance while minimizing imagination drift, exposing a WAM-specific vulnerability that undermines "check action against imagined future" safety claims.

view merged work →