ResearchPod Summary
Estimating 3D human motion from monocular video is essential for applications ranging from sports analytics and healthcare to robotics. However, current benchmarks often rely on expensive, laboratory-based motion capture (MOCAP) or constrained environments, which limits their ability to evaluate model performance in real-world, unscripted settings. The authors introduce CalTennis, a large-scale, multi-view video dataset designed to address this gap by providing a label-free, in-the-wild benchmark for evaluating monocular-to-3D pose estimation.
CalTennis comprises over 11 million frames of tennis practice and match play, captured using 2–6 synchronized consumer cameras. The authors leverage the standardized geometry of tennis courts to perform automated camera calibration and temporal synchronization, enabling a multi-view consistency framework. By comparing pose estimates across different camera angles, the authors can evaluate model accuracy without needing expensive ground-truth annotations. They also propose new metrics—specifically for footwork, stability, and body shape—to expose failure modes that standard benchmarks often overlook.
The authors benchmarked five state-of-the-art monocular 3D pose estimation models on the CalTennis dataset. While these models demonstrate high accuracy in recovering joint angles, they struggle significantly with metric-scale depth estimation, leading to unrealistic "jumps" in body position. Furthermore, the models exhibit inconsistent foot contact detection and varying body shape estimations (such as limb lengths) across different views. These findings suggest that while current pose estimation technology is advancing, it remains unreliable for downstream applications that require precise biomechanical or stability analysis.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.