ResearchPod Summary
High-fidelity free-viewpoint video and interactive rendering rely increasingly on explicit Gaussian representations. However, practical deployments are severely constrained by representation size, dynamic updates, and computational costs. Existing multi-view video benchmarks often use real-captured content that makes it difficult to isolate the effects of controlled camera geometry, representation efficiency, and temporal redundancy. To address these gaps, the authors introduce M3ISR, a richly annotated synthetic benchmark designed to evaluate both generative free-viewpoint video synthesis and efficient delivery.
M3ISR comprises 25 scenes divided across five indoor and outdoor categories, two camera-motion configurations, and six synchronized 1080p views. Every scene provides dense ground-truth annotations including RGB frames, calibrated camera parameters, depth, semantic and instance segmentation, and static-dynamic masks. The six cameras share a common scene center spanning a 120-degree horizontal fan, which intentionally eliminates translational parallax and isolates angular view variation. The benchmark is organized into five complementary tracks: 3DGS synthesis, 4DGS synthesis, 4DGS streaming, 3DGS compression, and 4DGS compression.
Representative baseline evaluations reveal clear practical trade-offs across the benchmark tracks. While evaluated static methods exhibit relatively small differences in reconstruction quality, they show substantial variation in representation storage requirements. Furthermore, evaluated streaming methods incur significantly higher training and reconstruction costs compared to offline dynamic reconstruction baselines. By standardizing camera geometry and ground-truth annotations, M3ISR establishes a rigorous testbed for future research into Gaussian-based representation efficiency.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.