ResearchPod Summary
This paper investigates whether the reward function used to train reinforcement learning (RL) agents for autonomous driving systematically shapes their internal attention mechanisms. While previous debates have questioned whether attention weights are faithful explanations of model reasoning, this study shifts the focus to a more fundamental question: does reward engineering predictably influence what an agent prioritizes in its visual input? The authors use three versions of a Perceiver-based driving agent, all sharing identical architectures and training data, but differing in their reward configurations (basic, minimal, and complete). They analyze cross-attention patterns across 50 real-world scenarios from the Waymo Open Motion Dataset (WOMD).
A key contribution of this work is the establishment of a robust methodology for analyzing attention in RL. The authors demonstrate that naively pooling timesteps across episodes significantly underestimates the relationship between collision risk and attention due to between-scenario heterogeneity. Instead, they propose using within-episode Spearman correlation aggregated via the Fisher z-transform. This approach reveals a clear, positive link between collision risk and agent-directed attention that is otherwise obscured by aggregate metrics.
The study identifies two primary ways reward design alters agent behavior. First, reward content directly shapes attention baselines: agents trained with navigation incentives allocate significantly more attention to GPS-path tokens compared to those without. Specifically, the minimal-reward model allocates 2.0x more attention to GPS tokens than the complete-reward model, and 4.7x more than the basic model. Second, the inclusion of continuous time-to-collision (TTC) penalties creates a 'learned vigilance prior.' Agents trained with these safety-critical rewards maintain elevated surveillance of other vehicles even during collision-free phases, demonstrating that the reward function can qualitatively shift an agent's attentional strategy rather than just modulating its intensity.
These findings suggest that attention analysis is a practical, diagnostic tool for developers of safety-critical RL systems. By inspecting where an agent allocates its representational resources, engineers can verify whether the reward function is successfully teaching the agent to monitor the most relevant aspects of the driving scene. This provides a path toward more transparent and reliable reward engineering in autonomous driving.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.