Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision

arXiv:2608.13283 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart Vanrumste, Benjamin Filtjens

KU Leuven · KU Leuven · KU Leuven · KU Leuven · Medical School Hamburg · KU Leuven · Delft University of Technology · Delft University of Technology

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted to ECCV Workshop 2026 (Human Motion Challenges in Real-World and Clinical Settings)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 50/100

The gist: This study presents the first use of egocentric vision for freezing of gait (FOG) detection in Parkinson's disease (PD) during home-based activities of daily living (ADLs).

Terminology

Summary

This study presents the first use of egocentric vision for freezing of gait (FOG) detection in Parkinson's disease (PD) during home-based activities of daily living (ADLs). The authors introduce egocentric video as a contextual modality for FOG detection, motivated by the fact that similar low-motion inertial patterns can arise from intentional stopping, object interaction, or balance recovery, making kinematic sensors alone insufficient for clinical motion understanding. The study recruited 15 individuals with PD (13 included in the final analysis) who completed two home-based data collection sessions (ON and OFF medication), each comprising five ADL tasks: walking through a doorway, two daily-life tasks randomly selected from a predefined list of 14, and two hotspot tasks where participants reported frequent FOG. Data were collected using five Xsens DOT IMUs (60 Hz) on the pelvis, bilateral shins, and feet, synchronized egocentric video from Pupil Core smart glasses (30 Hz, 1280×720), and four external cameras for offline expert annotation using ELAN software. The final dataset comprised 13 subjects and 171.6 minutes of recordings, with considerable inter-subject variability in FOG frequency and duration.

The experimental setup evaluated frozen representations from pretrained foundation models under leave-one-subject-out (LOSO) cross-validation. For IMU signals, the authors used UniMTS (a motion-specialized model pretrained for human activity recognition) and Chronos-2 (a general-purpose multivariate time-series model). For ego-video, they evaluated four models: DINOv3 (single-frame image model), VideoMAE-v2 (third-person video model encoding 16 frames), V-JEPA 2 (third-person video model encoding 64 frames), and EgoVideo (egocentric video model encoding 4 frames). All foundation model features were evaluated with a linear probe (L2-regularized logistic regression), and two fully trained temporal convolutional network (TCN) baselines were trained end-to-end on IMU signals (accelerometer-only and accelerometer+gyroscope variants). Windows were 2 seconds long with a 0.5-second stride, and a window was labeled FOG if it contained at least 0.5 seconds of annotated FOG (median episode duration was 0.9 seconds).

The main results at the 2-second window length show that the IMU-based TCN trained from scratch achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Among the IMU foundation models, Chronos-2 was the stronger representation (F1 38.7 vs. 29.1 for UniMTS) and operated at a much lower false positive rate (6.1% vs. 28.8%), whereas UniMTS achieved higher recall (56.7% vs. 39.9%). Adding gyroscope data to the TCN baseline lowered F1 relative to accelerometer alone (35.5 vs. 42.3) but reduced false positives (FPR 5.8% vs. 7.2%). Among ego-video models, V-JEPA2 was the strongest, surpassing the motion-specialized UniMTS on F1 (32.6 vs. 29.1), AUROC (77.2 vs. 73.5), and FPR (17.6% vs. 28.8%). The single-frame DINOv3 representation was weakest (F1 17.0, AUROC 53.6), and moving from one frame to sixteen frames (DINOv3 to VideoMAE-v2, which share the same ViT-B backbone and output dimensionality) raised AUROC from 53.6 to 68.6. All three video encoders significantly exceeded the single-frame baseline on AUROC (68.6, 72.4, and 77.2 vs. 53.6; p < 0.01), indicating that egocentric video carries FOG-relevant information beyond a single frame. Notably, capacity did not explain performance differences: the largest backbone (EgoVideo, ≈1 B parameters) did not outperform V-JEPA 2 (≈300 M parameters; F1 27.7 vs. 32.6).

A window-size ablation (2s, 3s, and 10s) showed that F1 generally increased with window length (e.g., Chronos-2 from 38.7 to 42.4, V-JEPA2 from 32.6 to 36.4), likely due to the higher proportion of FOG windows and reduced class imbalance, while AUROC consistently decreased (e.g., Chronos-2 from 82.9 to 71.6, V-JEPA2 from 77.2 to 66.4), suggesting that longer windows reduce separability due to more heterogeneous samples and noisier labels. By AUROC, Chronos-2 remained the strongest model at every window length, and V-JEPA2 was the best ego-video representation at the shorter 2s and 3s windows. A stride ablation (0.5s, 1s, 1.5s) showed performance remained broadly stable across strides for most models.

Qualitative analyses revealed that in several tasks, both modalities detected FOG episodes at similar time points (e.g., subjects 007, 010 Doorway), suggesting ego-vision contains independent predictive information. Some recordings showed the ego-vision model identifying prolonged stopping better than the IMU models, but this did not hold across the cohort. A direct test on annotated voluntary-stop windows (containing at least 0.5s of stopping and no FOG) showed that V-JEPA2 produced more false alarms than both Chronos-2 (24.9% vs. 9.7%) and the acc-only TCN (11.0%), and every model false-alarmed more during stopping than on other non-FOG windows. The authors conclude that voluntary stopping is therefore a failure mode shared across modalities, and whether visual context helps appears to vary by recording. The vision model also missed some FOG episodes correctly identified by the IMU models, and in a few cases ego-vision performed much worse than the IMU baseline (e.g., subjects 015 and 014). For subject 015, who had no annotated FOG episodes, V-JEPA2 generated many false positives, possibly because prolonged stopping occurred during object interaction (e.g., cleaning a cat litter box) while hands were not visible. Subject 014 displayed poor performance, potentially due to low-light recording conditions degrading visual representation quality.

The authors acknowledge several limitations: the relatively small cohort and short semi-structured ADLs; the reliance on frozen features from pretrained foundation models without fine-tuning (suggesting low-rank adaptation as future work); the lack of multimodal fusion between IMU and ego-vision; and the offline classification pipeline not being directly applicable to real-time use. They conclude that ego-vision alone performed competitively with IMU-based models, providing evidence that visual context contains independent cues relevant to FOG, but that a direct test on annotated stop windows, however, showed that ego-vision FM features alone do not disambiguate voluntary stopping from freezing better than inertial sensing. The findings motivate future work on context-aware multimodal approaches, including foundation model fine-tuning, fusion strategies, and real-time deployment for home-based clinical motion monitoring.

Improvements for AI systems

Improvements to AI systems:

  1. Multimodal fusion architecture – Combine IMU (Chronos-2) and ego-video (V-JEPA 2) features via cross-attention or late fusion with learned weighting, trained end-to-end. This exploits their complementary strengths (IMU: high recall; ego-vision: lower false positives on motion patterns) to improve F1 and reduce false alarms on voluntary stops.

  2. Context-aware false-positive suppression – Add a dedicated classifier head that explicitly distinguishes voluntary stopping from FOG using ego-video temporal context (e.g., object interaction, hand visibility, scene transitions). Train on annotated stop windows to penalize false alarms, reducing V-JEPA2's 24.9% false-alarm rate on stops toward Chronos-2's 9.7%.

  3. Fine-tuning with low-rank adaptation (LoRA) – Apply LoRA to both IMU and video foundation models (Chronos-2, V-JEPA 2) on the PD dataset instead of frozen features. This adapts representations to FOG-specific motion and visual patterns, likely improving AUROC beyond the reported 77.2–83.0 range.

  4. Adaptive window-length selection – Implement a dynamic windowing mechanism that uses short windows (2s) for high separability (AUROC) and longer windows (10s) for higher recall (F1), with a confidence-based switch. This balances the observed trade-off (F1 increases, AUROC decreases with window length).

  5. Lighting-robust video encoding – Preprocess ego-video frames with illumination normalization or use a vision model fine-tuned on low-light data, addressing subject 014's poor performance due to dark recording conditions. This improves robustness for home environments.

  6. Temporal attention for episode detection – Replace linear probe with a lightweight temporal convolutional or transformer head over V-JEPA2 features (64-frame encoding) to capture episode-level dynamics, improving detection of short FOG episodes (median 0.9s) beyond single-window classification.

  7. Real-time streaming pipeline – Convert the offline windowed classifier into an online system using a sliding window with 0.5s stride and a latency-bounded decision rule (e.g., trigger alarm after 2 consecutive FOG windows). This enables wearable+smartglass deployment for home monitoring.

What the improved AI system can do:

  • Detect FOG events during daily activities with higher F1 (target >50) and AUROC (>85) than any single modality, by fusing IMU and ego-vision.

  • Distinguish voluntary stopping from freezing with >80% accuracy, using visual context (e.g., object interaction) to suppress false alarms.

  • Operate robustly in low-light home conditions and across varied tasks (doorway, cleaning, etc.) without performance collapse.

  • Provide real-time alerts with <3s latency, enabling timely intervention during OFF-medication episodes.

  • Adapt per-user via fine-tuning, reducing inter-subject variability (e.g., handling subjects with no FOG episodes like 015).

Sources

Related papers