Author-updated Summary
Verified author edit
Recent advances in Multimodal Large Language Models have driven rapid progress in video understanding, yet existing benchmarks predominantly rely on web-sourced or movie videos that lack true inter-clip spatiotemporal continuity. Real-world daily life presents a fundamentally different challenge: long periods of redundant routines punctuated by sparse, decisive events separated by days or weeks. This paper introduces EgoMonth to investigate whether current video LLMs can maintain faithful, long-term spatiotemporal memory across month-scale egocentric experience, moving beyond lossy short-term clip summarization.
The EgoMonth dataset comprises over 300 hours of first-person daily-life video collected from 20 participants over spans of 20 to 120 days. To address privacy concerns, the authors developed an automated and manual anonymization pipeline using Grounding DINO and SAM 2 to blur faces, personal identifiers, and sensitive documents while preserving essential spatiotemporal context. The benchmark features 1,443 human-crafted multiple-choice question-answer pairs spanning single-video and cross-video queries. The evaluation framework organizes 14 tasks into three hierarchical cognitive levels:
Evaluating twelve representative open-source and closed-source MLLMs demonstrates a consistent performance decline along the hierarchical cognitive levels, with models performing best on Level 1 and worst on Level 3. Gemini 2.5 Pro achieves the highest overall macro-average accuracy at 71.8%, outperforming open-source models like Qwen2.5-VL (32B) at 58.0%, but still remains 22.4 percentage points below the human baseline of 94.2%. Furthermore, several models perform near or below the 25% chance level on demanding spatial and temporal tasks, demonstrating that simple increases in frame density or model parameters do not guarantee robust long-term memory.
EgoMonth establishes a rigorous evaluation standard for long-term embodied AI and lifelogging systems. By exposing fundamental limitations in current MLLM architectures—such as temporal attention dilution and spatial grounding collapse—the benchmark highlights the urgent need for novel video models capable of genuine, faithful, and multi-week spatiotemporal memory.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.