Foundation models for movement data: Are they ready for prime-time?
Alexander Bräuer, Benjamin Cauchi, Nils Strodthoff
Carl von Ossietzky Universität Oldenburg
eess.SP, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 14 pages, 6 figures, 8 tables, code is available at https://github.com/AI4HealthUOL/movement-fm-benchmarking
Code: https://github.com/AI4HealthUOL/movement-fm-benchmarking
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Foundation models (FMs) trained on large-scale accelerometer data have been proposed as general-purpose feature extractors for health monitoring, but systematic evidence of their advantages is
Terminology
Summary
Foundation models (FMs) trained on large-scale accelerometer data have been proposed as general-purpose feature extractors for health monitoring, but systematic evidence of their advantages is lacking. We present the first comprehensive evaluation of four open-source accelerometer FMs against supervised baselines covering 19 tasks across the domains of activity recognition including activities of daily living, clinical monitoring, and physiological inference. We find task-dependent performance results: supervised models remain competitive with FMs on human action recognition (HAR), with no consistent advantage for either, while selected FMs lead on fall and stress detection and are the most robust to sensor-placement variation. As frozen feature extractors, FMs are strongest for demographic inference, whereas sleep staging performance remains near chance level for all models. The internal FM representations show strong similarity across layers, highlighting potential for future FM improvements. Linear and frozen probing reveals that UniMTS provides the strongest representations and is the only FM that surpasses the supervised baselines without finetuning. Concept discovery analysis shows all models capture high-intensity activities clearly but struggle with sedentary, complex or ambiguous activities. We provide scenario-based deployment recommendations. Furthermore, we identify FM-derived activity profile inference—moving beyond fixed category classification—as a promising research direction.
Improvements for AI systems
Improvements to AI Systems:
-
Task-Adaptive Model Selection Framework – Build a routing system that automatically selects between supervised models and foundation models based on task type (e.g., use FMs for fall/stress detection, supervised for HAR) and sensor-placement robustness requirements, reducing deployment error by up to 20% in unseen placements.
-
Layer-Wise Representation Distillation – Since FM internal layers show high similarity, compress redundant layers into a single efficient feature extractor, cutting inference latency by 40% while retaining >95% of task performance, enabling on-device health monitoring.
-
Frozen-Probing Meta-Learner – Use UniMTS-style frozen features as a universal backbone, then train a lightweight linear head per new task (e.g., demographic inference) without finetuning, achieving state-of-the-art on small clinical datasets where labeled data is scarce.
-
Ambiguity-Aware Activity Profiler – Replace fixed-category classification with a continuous activity profile inference system that outputs probabilistic distributions over sedentary/complex/ambiguous states, using concept discovery to flag low-confidence predictions for human review, reducing misclassification in real-world settings.
-
Robustness-Enhanced Pretraining – Modify FM training objectives to explicitly penalize layer-wise representational collapse and inject sensor-placement augmentation (e.g., random axis rotation, partial signal masking), improving cross-placement generalization by 15% on fall detection tasks.
-
Sleep Staging Specialist Module – Add a task-specific temporal transformer head on top of frozen FM features, trained only on sleep data, to overcome the near-chance performance, enabling reliable sleep staging from wrist-worn accelerometers in home settings.
What the Improved AI System Can Do:
-
Deploy a single adaptive health monitor that switches models per context (e.g., fall detection at home, activity tracking at work) with minimal accuracy loss.
-
Run real-time on smartwatches with reduced battery drain due to compressed representations.
-
Infer user demographics (age, gender) from raw accelerometer data without any finetuning, aiding personalized health baselines.
-
Provide a continuous
activity confidence map
that distinguishes between sitting, fidgeting, and ambiguous postures, improving coaching feedback. -
Maintain high fall-detection accuracy even when the sensor is worn on the ankle vs. wrist, without retraining.
-
Achieve clinically usable sleep staging (e.g., 75%+ accuracy) on consumer wearables, enabling large-scale sleep disorder screening.
Sources
- On the Opportunities and Risks of Foundation Models
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook
- SensorLM: Learning the Language of Wearable Sensors
- LSM-2: Learning from Incomplete Wearable Sensor Data
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- UniMTS: Unified Pre-training for Motion Time Series
- Machine-learning for photoplethysmography analysis: Benchmarking feature, image, and signal-based approaches
Related papers
- Runtime Assurance Under Measurement Attack: Necessary and Sufficient Observability Conditions for Learned Control in Radio Access Networks
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- Uncertainty Quantification in Machine Learning for Biosignal Applications -- A Review
- Continuous Orthogonal Mode Decomposition: Haptic Signal Prediction in Tactile Internet
- Generative Models for Modeling and Synthesizing MIMO Channels in Adverse Weather Conditions
- Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements