Beyond Simulated Benchmarks: Evaluating Motion Representations for Fall Detection Under Real-World Data Scarcity

arXiv:2608.13197 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Timilehin B. Aderinola, Ilaria D'Ascanio, Luca Palmerini, Lorenzo Chiari, Jochen Klenk, Clemens Becker, Brian Caulfield, Georgiana Ifrim

University College Dublin · University of Bologna · Robert Bosch Hospital · Heidelberg University Hospital

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/mlgig/fall-lm

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention.

Terminology

Summary

Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. However, real-world falls are extremely rare: collecting 100 of them requires an estimated 100,000 days of monitoring, resulting in severely limited labelled data for training machine learning models. Consequently, many approaches rely on simulated datasets, often reporting high laboratory performance but limited real-world generalisation. We present a systematic evaluation of motion representations for wearable fall detection under real-world data scarcity. Using accelerometer signals, we compare interval-based, kernel-based, symbolic, and foundation model representations. As an interpretable baseline, we additionally investigate a lightweight symbolic representation that converts short motion segments into symbolic sentences augmented with physically-grounded impact descriptors. Experiments use FallAllD, a simulated falls dataset, and FARSEEING, a clinically verified real-world falls dataset. Through cross-validation, controlled data scarcity, and cross-dataset transfer, we examine how representation choices affect robustness under realistic deployment. Our results reveal that highly parameterised kernel and foundation models excel on simulated data but degrade severely under both data scarcity and domain shift. Although the interval-based representation achieves the strongest absolute real-world performance, augmenting a symbolic representation with physically-grounded impact descriptors yields the smallest degradation under domain shift and retains detection sensitivity under extreme scarcity, albeit at lower precision. These findings highlight the importance of evaluating beyond simulated benchmarks and show that representation choice is critical for deployable fall detection given the scarcity of real-world data.

The main contributions of this work are as follows:

  1. We present a systematic comparison of interval-based, kernel-based, symbolic, and foundation-model representations for wearable fall detection, using a unified streaming evaluation pipeline with subject-wise splits.

  2. We characterise representation robustness under extreme data scarcity and quantify the simulation-to-reality gap through cross-dataset transfer from simulated falls to clinically verified real-world falls.

  3. As an interpretable case study, we implement FallLM, a lightweight symbolic representation augmenting SAX tokens with impact descriptors, and use its transparency to diagnose failure modes at the token level.

  4. We release our code to support reproducible benchmarking of wearable fall detection and related time series event detection tasks.

Our experiments use accelerometer data from the FARSEEING real-world falls dataset and the publicly available FallAllD dataset, which contains simulated falls performed in controlled laboratory settings. To ensure realistic evaluation, we adopt a streaming event-detection protocol with subject-wise splits. During training, samples are generated using fixed-size overlapping windows extracted from the continuous signals. During testing, signals are processed in a continuous streaming setting, beginning at arbitrary points independent of the fall impact event. For each incoming window, the trained model estimates the probability of a fall, which is then thresholded to produce fall predictions.

In this work, using a realistic streaming evaluation pipeline, we systematically evaluate motion representations for wearable fall detection. We compare interval-based, kernel-based, symbolic, and foundation-model representations under realistic constraints of limited real-world fall data and cross-dataset transfer between simulated and real-world datasets, with particular attention to how representations behave when moving from simulated to real-world falls.

The tri-axial acceleration signals are converted into a single acceleration magnitude signal, computed as M = sqrt(Accx squared + Accy squared + Accz 2), where Accx, Accy, and Accz are the anterior–posterior, medial–lateral, and vertical components. Using magnitude reduces sensitivity to sensor orientation and enables consistent processing across devices and placements. Fall samples are extracted using three-second windows beginning one second before the annotated impact, comprising a one-second pre-impact phase, the impact, and a one-second post-impact phase. This compact window captures the essential dynamics of the fall while remaining suitable for real-time detection. ADL (negative) samples are extracted from fall-free portions using one-second-step sliding windows, retaining only windows whose maximum magnitude exceeds 1.4 g; this removes low-intensity movement while preserving dynamic activities such as walking or turning.

Since FARSEEING only provides binary impact labels for each timepoint, we pose fall detection as binary classification (fall vs. ADL). To simulate real-world deployment, fall detection is evaluated in a streaming setting where the continuous accelerometer signal is processed sequentially without prior knowledge of fall events, following a recently proposed streaming evaluation protocol. The signal is analysed using sliding windows with a one-second step size. For each window, the trained model estimates the probability of a fall event. To account for the temporal context of fall events, detection is evaluated within a tolerance interval surrounding the annotated impact point. This interval includes both the motion leading to the fall and the immediate recovery period following impact. A detection is considered correct if the predicted fall window overlaps with the ground-truth interval. Overlap between predicted windows and ground truth is measured using the Intersection over Union, defined as IoU(d, R) = d ∩ R/d ∪ R, where d denotes the detected window and R the ground-truth interval. A detection with IoU(d, R) > 0 is counted as a true positive, while detections without overlap are considered false positives. If no detection overlaps with the fall interval, the event is counted as a false negative. To avoid registering repeated firings of the sliding window as separate detections, a debounce step retains only the first alarm within a detection episode.

Each evaluated representation transforms the acceleration magnitude into a feature space that a classifier then uses. Symbolic representations transform continuous signals into sequences of discrete symbols drawn from a finite alphabet. This discretization process reduces sensitivity to noise and allows temporal patterns to be represented as symbolic words that can be analysed using techniques similar to those used in text processing. In this work, we evaluate three symbolic approaches that employ different symbolic transformations and modelling strategies: WEASEL, MrSQM, and FallLM.

WEASEL 2.0, based on the Symbolic Fourier Approximation (SFA), applies a random selection of dilated windows, extracting SFA words at multiple scales and dilations to form a high-dimensional dictionary. The resulting feature vector is classified with a ridge regression classifier. MrSQM discretises segments of the time series into symbolic tokens (SAX or SFA) and constructs features from symbolic subsequences extracted across multiple resolutions, which are then classified using logistic regression.

FallLM is a lightweight symbolic representation for motion signals. The acceleration magnitude signal is first segmented into short intervals using Piecewise Aggregate Approximation (PAA), after which the aggregated values are discretised into symbolic tokens using SAX. Each window is therefore represented as a short sequence of symbols capturing coarse motion dynamics. SAX discretization uses Gaussian breakpoints derived from the standard normal distribution, yielding a standardised symbolic alphabet. In addition to the SAX token sequence, we augment each motion sentence with two impact-related tokens derived from the peak acceleration magnitude within the window. The first token encodes the absolute impact level by discretising the peak acceleration into three levels informed by the biomechanical ranges reported in [10]: low (< 1.8g, typical for normal walking), medium (1.8g–2.5g, associated with deliberate vertical displacement such as climbing stairs), and high (> 2.5g, associated with hard impacts or falls). Because these thresholds are defined in units of gravitational acceleration (g), they remain consistent across datasets and sensor configurations. We show that this coarse, dataset-invariant encoding transfers better than an absent or finer-grained token. To capture contextual information about the impact relative to surrounding motion, we also include a relative impact token that reflects the peak acceleration relative to the magnitude distribution within the window. This token provides a coarse indication of how abrupt the impact is compared to the surrounding movement dynamics. Together, the absolute and relative impact tokens complement the SAX representation by incorporating information about motion intensity while preserving the symbolic structure of the motion sequence. The resulting symbolic sequence is converted into an n-gram bag-of-words representation using a TF-IDF (Term Frequency-Inverse Document Frequency) vectoriser, and a logistic regression classifier is trained on these features for fall detection.

MiniRocket transforms the input signal using a large number of fixed convolutional kernels and extracts simple statistics from the resulting feature maps, such as the proportion of positive values. These features form a high-dimensional representation that can be efficiently classified using a linear model (Ridge Regression Classifier). QUANT extracts quantile statistics from hierarchical intervals of the time series, capturing distributional properties of the signal at different temporal locations. These interval features are then used by an ensemble classifier (ExtraTrees Classifier). Mantis is a transformer-based model pre-trained using contrastive learning to capture general temporal patterns across datasets. In our experiments, the pre-trained model is finetuned to generate embeddings for fall detection.

For FallLM, the acceleration magnitude signal is discretised using SAX with an alphabet size of 5 and a word size of 3, fixed a priori from the temporal structure of fall events rather than tuned. A word size of 3 maps the 3-second window to the three biomechanical phases of a fall—pre-impact, impact, and post-impact (1 second per symbol)—and keeps motion sentences short enough to remain human-readable for the interpretability analysis. Each SAX symbol denotes an amplitude level from a (lowest) to e (highest). After appending the two impact tokens, each motion sentence comprises five tokens, vectorised as n-grams (range (4, 5)) and classified by logistic regression; this n-gram range forces the classifier to weigh the structural shape of the motion together with its impact magnitude.

We conduct three experiments: (1) subject-wise cross-validation on FARSEEING, (2) a data-scarcity analysis that progressively reduces the number of training falls, and (3) a cross-dataset transfer experiment from simulated (FallAllD) to real-world (FARSEEING) falls.

Each three-second window (300 samples) is labelled with a binary target indicating the presence or absence of a fall event. The test set consists of unsegmented signals, each containing a single fall annotated with a ground-truth impact index f. We define a tolerance interval of [f − 3, f + 20) seconds around the impact. The 3-second pre-impact allowance captures the falling phase, and the post-impact interval covers the period during which the person may remain on the ground. A detection window overlapping this interval is counted as a true positive. Given the extreme class imbalance of streaming detection, we report Precision, Recall, and F1 Score rather than accuracy. We additionally report Detection Delay (in seconds) and inference runtime per window as a measure of computational efficiency. Detection Delay is the signed time difference between the detection and the impact index; negative values indicate detection prior to the annotated impact, and smaller absolute values indicate more temporally precise detection.

The five-fold subject-wise cross-validation results on the FARSEEING dataset (an average of 31 falls per fold) show that the foundation model, Mantis, obtains the highest F1 score (0.83), with Quant, WEASEL, and MiniRocket close behind. Mantis achieves the highest recall (0.89) and Quant the highest precision (0.84), while WEASEL and MiniRocket maintain a balance between the two. These methods cluster tightly, indicating that with enough labelled real-world falls, interval-based, kernel-based, foundation model, and SFA-based symbolic representations all reach comparable in-domain performance. The two SAX-based symbolic methods, FallLM and MrSQM, achieve lower F1 scores (0.64 and 0.66). FallLM attains high recall (0.87), but the lowest precision (0.52). In a streaming setting with many ADL windows, this manifests as frequent false alarms. FallLM is the most computationally efficient method (0.01 ms per sample), though all evaluated methods are already fast enough for real-time and always-on deployment. Overall, these results show that several representation families achieve strong real-world performance when sufficient labelled data is available, and that the highest in-domain F1 does not single out one representational paradigm.

To examine behaviour under limited real-world data, we train each method on fractions (1%, 5%, 10%, 25%, 50%, 100%) of the 112 fall events in the training set while retaining all negative samples, averaging five runs over randomly sampled fall subsets per fraction. We run this on FARSEEING only, as simulated falls are not scarce. With only 1% of the available fall events (two falls), FallLM is the only method to achieve substantial detection (mean F1 ≈ 0.35), recovering falls with minimal supervision but at low and highly variable precision, depending on whether the two sampled falls are representative. Conversely, Mantis fires rarely but almost always correctly (near-perfect precision, low recall, mean F1 ≈ 0.11). This conservative behaviour may reflect the priors it inherits from pretraining. QUANT and MiniRocket detect only isolated falls, while MrSQM and WEASEL fail entirely (F1 = 0). Notably, the methods that perform best with abundant data collapse under extreme scarcity, underscoring that strong in-domain performance does not imply robustness to scarcity. As more fall events become available, all methods improve rapidly, reaching near-peak performance by roughly 25–50% of available falls; several show slight precision declines thereafter, likely reflecting the heterogeneity of real falls and their similarity to fall-like ADLs. FallLM, however, plateaus near F1 = 0.64: its recall stays high, but its precision does not improve with additional data, so other methods overtake it on F1 once they accumulate enough falls. Among the symbolic methods, MrSQM does not detect falls until 25% of falls are available, suggesting that FallLM's scarcity-detection capability stems from its impact tokens rather than from symbolic representation in general. These results show that the choice of motion representation is decisive under data scarcity. Representations encoding strong, dataset-invariant priors can be particularly valuable when labelled falls are scarce. More broadly, effective fall detection may not require large collections of real-world falls, provided the representation is matched to the scarcity regime.

To assess robustness under domain shift, we train each method on the simulated FallAllD dataset and evaluate it on real-world FARSEEING falls without re-training or adaptation. To enable a direct comparison between in-domain and cross-domain performance, the same FallAllD-trained model is evaluated both on held-out simulated data and on the FARSEEING test folds. The simulation–reality gap is substantial and affects all methods. Mantis, MiniRocket, and WEASEL exceed F1 = 0.95 on simulated falls but fall to 0.37–0.50 on real falls, confirming that high simulated performance can badly overestimate real-world effectiveness. QUANT transfers best among the established methods (real F1 0.62), while FallLM achieves the highest real-world F1 (0.67) and the smallest relative degradation overall (9.5% drop in F1). The precision and recall values further reveal that the methods fail in distinct ways. MiniRocket, WEASEL, Mantis, and QUANT retain high precision under transfer but suffer large recall drops, the largest for WEASEL (-0.71). Having learned the morphology of simulated falls, they fire only on the real falls that most resemble them, missing the majority. The two SAX-based methods, FallLM and MrSQM, retain more recall under transfer (drops of 0.30 and 0.31, versus 0.46–0.71 for the others), indicating that the symbolic discretisation helps preserve sensitivity across domains. However, unlike MrSQM, which suffers a precision drop (-0.49), FallLM is the only method whose precision improves on real data (from 0.59 to 0.68). We attribute this to the absolute impact magnitude, expressed in physical units of g, which remains valid across datasets. This suggests that while symbolic discretisation aids recall retention, it is the physically-grounded impact tokens that additionally preserve precision. These results show that in-domain accuracy is a poor predictor of cross-dataset robustness, and that the type of representational prior matters more than raw performance: representations anchored to dataset-invariant physical quantities transfer better than those that model dataset-specific signal statistics.

A key advantage of FallLM's symbolic representation is that the learned model is inspectable. Each feature is a human-readable motion motif, and its logistic regression weight indicates how strongly it pushes a window toward the fall or ADL class. Fall and ADL motifs frequently share amplitude-shape prefixes and differ mainly in their impact token. For example, b d c impact med is a strong fall indicator (weight +1.83), while b d c impact low, identical in shape, pushes towards the ADL class (weight −0.98, not among the top motifs shown). Fall motifs are characterised by impact high and impact med (often with rel impact high), whereas ADL motifs are dominated by impact low. This confirms that FallLM's decisions rest on physically-grounded impact magnitude, and explains its limited in-domain precision: the boundary between falls and ADLs lies in the medium-impact regime, where vigorous activities and genuine falls are least separable. The acceleration waveforms underlying these motifs confirm that the impact tokens correspond to physically meaningful peaks. On FallAllD, every top fall motif carries impact high, so the model learns an exclusively high-impact notion of a fall. On FARSEEING, the fall vocabulary extends into the medium-impact range, reflecting that real-world falls span a wider range of intensities. This explains the transfer behaviour: a model trained on FallAllD inherits a high-impact-only fall vocabulary that transfers as a strict, high-precision detector on FARSEEING, where high impact is unambiguous, but which misses lower-impact real falls, capping its recall.

Our results point to a single overarching lesson: in-domain accuracy is a poor predictor of real-world robustness. The representations that dominate the simulated benchmarks degrade severely under domain shift. A practitioner selecting a method on simulated performance alone would choose among the least deployable options. This reframes how wearable fall detection should be evaluated: cross-dataset and scarcity behaviour are not secondary robustness checks but primary selection criteria. These results suggest that what governs transfer is not model simplicity but the type of prior a representation encodes. Methods that learn dataset-specific signal statistics capture the morphology of simulated falls and lose recall when real falls differ in shape; FallLM instead relies on absolute impact magnitude (in units of g), a prior whose meaning is invariant across datasets. This is supported by the recall decomposition and by MrSQM, a second SAX method that, lacking impact anchoring, collapses in precision whereas FallLM's improves. This robustness has a cost: FallLM is a probe of the principle rather than a deployable detector. Its anchoring makes it imprecise in-domain, where genuine falls and vigorous activities overlap in the medium-impact regime, and biased toward high-impact events. Trained on simulated falls, whose discriminative signal lies at the high-impact end, it inherits a high-impact-only notion of a fall and systematically misses lower-impact real falls. The same prior that confers robustness thus bounds the method's ceiling: a deployable detector would need to combine impact-anchored robustness with sensitivity to low-impact events.

For practice, these findings carry several implications. Cross-dataset evaluation should be treated as a standard reporting requirement rather than an optional analysis, since simulated performance can greatly overstate real-world effectiveness. Additionally, under data scarcity, representation choice matters more than model capacity. On the choice of detection paradigm, our binary formulation does not exploit the structure of normal activity that one-class methods model. However, such approaches face a specific difficulty in this domain: high-impact ADLs such as heavy sitting or jumping are statistically anomalous yet are not falls, so detectors keyed on unusual motion would tend to flag them. A systematic comparison of detection paradigms is left to future work.

Several limitations bound these conclusions. Our analysis rests on a single clinically-verified real-world dataset, and the scarcity of real fall data that motivates this work also limits the breadth of our cross-dataset claims. We evaluate one representative method per family, each with its canonical classifier, so observed differences reflect representation–classifier pipelines rather than representations in isolation. We deliberately study zero-shot transfer to measure the unadapted simulation–reality gap; domain adaptation would likely narrow it but would obscure the quantity we set out to characterise. Finally, the interpretability analysis reflects the highest-weighted motifs of a single model and is illustrative rather than exhaustive.

We presented a systematic evaluation of motion representations for wearable fall detection under real-world data scarcity and domain shift, comparing interval-based, kernel-based, symbolic, and foundation model approaches on simulated and clinically-verified real-world falls. Across cross-validation, controlled data scarcity, and cross-dataset transfer, we found that the methods strongest on simulated data degrade most under domain shift: the benchmark-topping model is not the most deployable one. Robustness instead tracks the type of prior a representation encodes. Representations anchored to dataset-invariant physical quantities transfer more gracefully than those modelling dataset-specific signal statistics. As an interpretable case study, FallLM illustrates both sides of this principle: its physically-grounded impact tokens yield the smallest transfer degradation, the highest real-world F1, and detection from as few as two fall examples, but at the cost of in-domain precision and a structural blindness to low-impact falls. As real-world fall data remains extremely scarce, we argue that cross-dataset evaluation should become standard practice, and that reconciling transfer robustness with precision is a key direction for future work. To support reliable benchmarking, we release our code and streaming evaluation pipeline, and continue to work toward the curation of open real-world fall datasets.

Improvements for AI systems

Improvements to AI Systems:

  1. Domain-Shift-Aware Representation Selection: AI systems should prioritize representations anchored to dataset-invariant physical quantities (e.g., absolute acceleration in g-units) over those modeling dataset-specific statistics, when deployment environments differ from training data. This improves cross-domain generalization.

  2. Scarcity-Robust Feature Engineering: Incorporate physically-grounded impact descriptors (e.g., absolute and relative peak acceleration tokens) into symbolic representations. This enables fall detection from as few as 2 training examples, whereas high-capacity models fail entirely under extreme data scarcity.

  3. Interpretable Failure-Mode Diagnosis: Use human-readable symbolic motifs (e.g., SAX tokens with impact labels) to enable token-level inspection of model decisions. This allows AI systems to identify systematic biases (e.g., high-impact-only fall vocabulary) and adjust thresholds or augment training data accordingly.

  4. Unified Streaming Evaluation Protocol: Implement subject-wise splits, continuous sliding-window inference, tolerance-based event matching, and debouncing to simulate real deployment. This prevents overestimating performance from segmented, impact-aligned test windows.

  5. Cross-Dataset Transfer as a Primary Metric: Require zero-shot cross-dataset evaluation (e.g., simulated-to-real) alongside in-domain metrics. This reveals that high simulated F1 (0.95+) can collapse to 0.37–0.50 on real data, guiding model selection toward robust options.

  6. Hybrid Representation for Precision-Recall Balance: Combine impact-anchored symbolic features (for recall retention under domain shift) with interval-based or kernel features (for precision in high-data regimes). This addresses FallLM's low in-domain precision while retaining its transfer robustness.

  7. Adaptive Detection Thresholding: Leverage the finding that precision degrades with more heterogeneous real falls. AI systems can dynamically adjust decision thresholds based on impact-level distributions (e.g., medium-impact ambiguity) to reduce false alarms without sacrificing recall.

  8. Physical-Unit Normalization: Encode features in physical units (g) rather than dataset-specific scales. This ensures consistency across sensor placements, devices, and populations, improving transferability without retraining.

What the Improved AI System Can Do:

  • Detect falls from real-world data with minimal labeled examples (as few as 2 falls), maintaining meaningful recall where current models fail.

  • Transfer from simulated to real-world settings with only 10% F1 degradation, versus 40–60% for benchmark-optimal models.

  • Provide transparent, token-level explanations for each detection, enabling clinicians to audit why a fall was flagged (e.g., high-impact vs. medium-impact event).

  • Avoid overconfident deployment decisions by reporting cross-dataset and scarcity performance as standard outputs, not optional analyses.

  • Maintain real-time inference (<0.01 ms per window) while achieving robust detection, suitable for always-on wearable devices.

  • Distinguish falls from vigorous ADLs (e.g., jumping, heavy sitting) by explicitly modeling impact magnitude relative to surrounding motion, reducing false alarms in medium-impact regimes.

Abstract

Falls are a major health concern for older adults, and wearable sensors have been widely explored for detecting falls and enabling timely intervention. However, real-world falls are extremely rare: collecting 100 of them requires an estimated 100,000 days of monitoring, resulting in severely limited labelled data for training machine learning models. Consequently, many approaches rely on simulated datasets, often reporting high laboratory performance but limited real-world generalisation. We present a systematic evaluation of motion representations for wearable fall detection under real-world data scarcity. Using accelerometer signals, we compare interval-based, kernel-based, symbolic, and foundation model representations. As an interpretable baseline, we additionally investigate a lightweight symbolic representation that converts short motion segments into symbolic sentences augmented with physically-grounded impact descriptors. Experiments use FallAllD, a simulated falls dataset, and FARSEEING, a clinically verified real-world falls dataset. Through cross-validation, controlled data scarcity, and cross-dataset transfer, we examine how representation choices affect robustness under realistic deployment. Our results reveal that highly parameterised kernel and foundation models excel on simulated data but degrade severely under both data scarcity and domain shift. Although the interval-based representation achieves the strongest absolute real-world performance, augmenting a symbolic representation with physically-grounded impact descriptors yields the smallest degradation under domain shift and retains detection sensitivity under extreme scarcity, albeit at lower precision. These findings highlight the importance of evaluating beyond simulated benchmarks and show that representation choice is critical for deployable fall detection given the scarcity of real-world data.

Sources

Related papers