ResearchPod Summary
How can we comprehensively evaluate and improve the ability of multimodal large language models to understand professional basketball games, which requires integrating visual perception, spatial-temporal localization, and structured domain knowledge?
The authors introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten diverse tasks using text, image, and video data from the 2025-2026 NBA season. To address the performance gaps revealed by the benchmark, they propose BasketballSkills, an agent framework that uses a language-model controller to dynamically select and coordinate eight atomic perception and retrieval tools under four reusable procedural skills.
Experiments demonstrate that existing commercial and open-source multimodal models struggle significantly when questions require the combination of multiple reasoning and perception capabilities. In contrast, the BasketballSkills framework outperforms the strongest commercial multimodal model on eight out of the ten benchmark tasks. This indicates that explicitly structuring domain-specific capabilities into reusable, verifiable workflows is highly effective for complex sports understanding.
This work provides a rigorous evaluation standard and a modular architecture for sports video understanding. By moving beyond isolated task evaluations, it establishes a foundation for building knowledge-grounded AI systems capable of professional-level game analysis.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.