ResearchPod Summary
Traditional Referring Expression Comprehension (REC) in UAV imagery is limited by its reliance on text-only queries and the assumption that a single target exists in the scene. This paper addresses the need for a more flexible framework—termed 'Universal Referring'—that can handle diverse query modalities (text, images, or both) and variable target counts (zero, one, or multiple) in complex, small-object-heavy aerial environments.
The authors construct the UniRef-UAV benchmark, which synthesizes data from 22 public UAV datasets. It includes over 150,000 query samples across three modalities: text-only, image-only, and text+image. The benchmark supports both in-domain and cross-domain evaluation protocols. To establish a baseline, the authors propose UAV-URNet, a detection-style model that maps heterogeneous queries into a shared latent space. By utilizing a set-prediction framework, the model can output a variable number of bounding boxes, allowing it to handle 'no-target' scenarios and multi-instance grounding without requiring architectural changes for different query types.
Experiments demonstrate that existing models trained on ground-level datasets (like gRefCOCO) or standard text-only UAV datasets perform poorly on the UniRef-UAV benchmark, highlighting the necessity of this new, specialized dataset. The UAV-URNet baseline provides a stable and reproducible performance, showing that multimodal queries (combining text and images) significantly reduce visual ambiguity compared to text-only or image-only inputs. The results suggest that unifying query-target alignment in a shared space is a robust strategy for complex aerial scene understanding.
As UAVs are increasingly deployed for tasks like search and rescue or surveillance, they must interpret complex, multi-modal instructions to identify targets accurately. By moving beyond the 'single-text-to-single-object' paradigm, this work provides the necessary tools and benchmarks to develop more versatile, intelligent aerial vision systems capable of handling real-world ambiguity.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.