ResearchPod Summary
Understanding clinical motion in daily living environments is an important challenge for wearable and computer vision systems. Unlike structured laboratory assessments, daily-life activities performed at home are highly variable and shaped by the surrounding environment and ongoing tasks. Commonly used kinematic sensors may be insufficient for determining clinical motion, as similar low-motion inertial patterns can arise from intentional stopping, object interaction, or balance recovery. This ambiguity is particularly problematic for freezing of gait (FOG) in Parkinson's disease, an episodic symptom strongly influenced by contextual triggers such as narrow doorways, turning, or dual-tasking. While automatic FOG detection has traditionally relied on physiological and inertial measurement units (IMUs), distinguishing FOG from voluntary stopping remains difficult. This paper investigates whether egocentric vision can provide the necessary task-related visual context to improve clinical motion understanding and FOG detection in home environments.
The authors collected a synchronized home-based dataset from 13 Parkinson's disease participants during activities of daily living across ON and OFF medication states. The sensor setup included five Xsens DOT IMUs placed on the pelvis, shins, and feet, along with Pupil Core smart glasses capturing egocentric video at 30 Hz. Expert annotations following the updated FOG video scoring definition served as ground truth. Data were segmented into 2-second windows, and representations were extracted from a suite of pretrained foundation models for both modalities—including UniMTS and Chronos-2 for IMUs, and DINOv3, VideoMAE-v2, V-JEPA 2, and EgoVideo for egocentric video. All models were evaluated under a leave-one-subject-out cross-validation protocol using linear probing, alongside a fully supervised IMU-based Temporal Convolutional Network (TCN) trained from scratch as a reference baseline.
The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC. In contrast, the best-performing egocentric video model, V-JEPA 2, attained 32.6 F1 and 77.2 AUROC. Although egocentric video representations alone did not outperform wearable inertial sensing, they demonstrated clear above-chance discrimination. Qualitative analyses of modality-specific errors further indicated that egocentric vision captures FOG-relevant contextual cues independently of IMUs, suggesting that combining first-person visual context with inertial data could yield more robust motion understanding systems for daily living.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.