ResearchPod Summary
Traditional semantic segmentation assigns pixels to a closed set of object classes, which fails when autonomous vehicles encounter rare or unknown obstacles in the real world. Naming an unknown object is insufficient for motion planning; the system must understand whether a region is drivable and how severe a collision would be. This paper investigates whether dense action-relevant attributes, specifically 7-rank drivability and 5-rank vulnerability, can be predicted directly from the hidden states of a vision-language model.
The proposed method, VOLA, feeds an input image and a short prompt into a VLM (Qwen3.5), tapping image-token hidden states from an intermediate layer to form a coarse spatial semantic grid. A lightweight boundary-aware decoder upsamples this grid to full resolution while incorporating RGB appearance features from a MobileViT-XXS branch. Unlike prior VLM segmentation approaches, VOLA requires neither autoregressive text generation nor external mask models like Segment Anything (SAM).
Because standard real-world datasets lack comprehensive lane connectivity and out-of-distribution attribute labeling, the authors construct a dense supervision dataset using the CARLA simulator. Autopilot driving data is collected across varied weather conditions and filtered to remove idle frames. Per-pixel labels are derived from three sources: lane topology mapped from simulator waypoints, temporary traffic-light availability, and object semantics from a segmentation camera. To ensure strict evaluation without data leakage, the dataset is split by town, using four towns for training and held-out towns for validation and testing.
Experiments test VOLA against vision-only segmenters and prompted VLM segmenters under progressively harder distribution shifts. VOLA matches strong vision-only segmenters on familiar categories while significantly improving transfer to real open-world anomalies. On anomalous objects, VOLA achieves 69.4% mean vulnerability-rank recall, compared to 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results demonstrate that VLM image tokens provide powerful semantic cues that transfer effectively to unseen obstacles outside the training vocabulary.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.