ResearchPod Summary
Sekai2 is a multi-source, real-world video dataset specifically curated to advance the development of interactive world models. Unlike existing large-scale video corpora that prioritize visual diversity without spatial grounding, or reconstruction-oriented datasets that focus on short-range geometry, Sekai2 provides long-horizon temporal continuity. It includes 128,892 clips across 113 countries, with a significant portion of the data consisting of continuous, two-minute segments. This structure allows models to learn how to maintain persistent scene states and handle camera control over extended durations.
The authors constructed Sekai2 by integrating three distinct data sources: a re-processed version of the original Sekai dataset, newly collected YouTube footage (including walking tours, driving, and drone shots), and a specialized set of 982 panoramic sequences. A unified data engine was developed to handle shot-aware clip construction, quality filtering, and camera-pose estimation. To address the challenge of long-term spatial memory, the panoramic subset features non-linear trajectories with loops and revisits, which were refined using geometrically verified loop closures to minimize trajectory drift.
A key contribution of Sekai2 is its hierarchical semantic annotation schema. Each clip is annotated at both the global and interval levels, disentangling subject motion, environmental dynamics, static scene content, and camera behavior. By providing these labels on a shared timeline, the dataset enables models to explicitly distinguish between changes in the world and changes in the observer's viewpoint. The inclusion of both full and short prompts—where camera-specific information can be toggled—further supports research into camera-controllable video generation.
As video generation shifts from simple synthesis to interactive simulation, the ability to maintain spatial consistency and respond to user or camera input becomes paramount. Sekai2 provides the necessary supervision to train models that can navigate, remember, and interact with complex, real-world environments. By offering a large-scale, consistent, and temporally grounded resource, it addresses a critical bottleneck in the current research landscape for long-horizon world modeling.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.