ResearchPod Summary
As robotic systems increasingly adopt World-Action Models (WAMs)—which couple action generation with future world prediction—a critical security question arises: can these models be manipulated to fail while maintaining the appearance of safe, coherent behavior? The authors investigate whether the internal alignment between a robot's 'imagination' and its 'execution' is robust against adversarial visual perturbations.
The researchers introduce BadWAM, a unified framework for executing 'World-Action Drift Attacks.' The framework operates by injecting small, bounded visual perturbations into the robot's input stream at inference time. It explores two primary attack strategies:
The authors evaluate these attacks on closed-loop robotic manipulation tasks using various WAM architectures, treating the models as black-box systems where the attacker can query outputs but does not have access to internal gradients.
The study demonstrates that WAMs are significantly more fragile than previously assumed. In action-only settings, the authors observed a drastic reduction in task success rates, with one model dropping from 96.5% to 43.1% performance. More importantly, the imagination-preserving attacks reveal a specific vulnerability: it is possible to desynchronize action and imagination. This means a robot can be forced to perform incorrect, task-failing actions while its internal 'world model' continues to predict a plausible, successful future. This creates a dangerous gap, as safety monitors that rely on inspecting imagined futures may fail to detect the ongoing attack.
This research highlights a fundamental flaw in the 'imagine-then-act' paradigm for embodied AI. If safety mechanisms rely on the assumption that a plausible imagined future implies a safe action, they are inherently vulnerable to drift attacks. The findings suggest that future safety protocols must verify the alignment between predicted outcomes and actual control commands, rather than trusting the imagination in isolation.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.