ResearchPod Summary
SLVMBench is a new benchmark designed to evaluate how well video large language models (video-LLMs) can acquire procedural skills from long-form video memory and apply them to real-time tasks. Unlike existing benchmarks that focus on passive video comprehension or short-term demonstrations, SLVMBench requires models to process a tutorial video, maintain that knowledge while navigating through hours of irrelevant distractor videos, and then correctly apply that knowledge to a target task at a precise temporal cutoff.
The researchers curated a dataset of 2,261 human-generated question-answer pairs across 11 categories, ensuring all videos were published after January 2025 to prevent data contamination. Each evaluation instance consists of a tutorial video, a sequence of distractor videos, and a target video. The benchmark tests three settings: a baseline without a tutorial, an immediate application setting, and a long-memory setting where the tutorial is separated from the target task by up to two hours of video. The tasks are designed to be non-trivial, requiring models to perform procedural mastery, understand tool logic, and handle diagnostic scenarios.
The study reveals a significant "performance cliff" for state-of-the-art video-LLMs. While models consistently improve when provided with immediate tutorial information, their performance drops drastically when that information must be retrieved from long-term memory after a two-hour gap. In many cases, the performance gain from the tutorial nearly vanishes, suggesting that current architectures lack the robust episodic memory required for real-world agentic tasks.
As AI agents move toward real-world applications—such as assisting with software operations or physical repairs—they must be able to learn from long-form demonstrations and recall that information later. SLVMBench provides a rigorous, high-fidelity testbed for researchers to identify the limitations of current memory mechanisms and develop models capable of true long-horizon procedural reasoning.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.