Mobile Interaction for Assessing Fatigue, Sleep, and Activity in Neurodegenerative and Chronic Diseases

arXiv:2608.06380 · cs.HC, cs.AI · Submitted 2026-06-09 · Read on arXiv

Julian Fierrez, Alejandro Peña, Aythami Morales, Ruben Tolosana, Ruben Vera-Rodriguez, Meenakshi Chatterjee, Ahmaniemi Teemu, Wan-Fai Ng, Walter Maetzler, Nikolay V. Manyakov, Jennifer Kudelka, Ralf Reilmann, C. Janneke van der Woude, Kristen Davies, Victoria Macrae

Universidad Autonoma de Madrid · Johnson & Johnson · VTT Technical Research Centre of Finland · Newcastle upon Tyne Hospitals NHS Foundation Trust · University Hospital Schleswig-Holstein and Kiel University · George Huntington Institute · Erasmus University Medical Centre

cs.HC, cs.AI

Submitted: 2026-06-09

Comments: IEEE Conf. on Computers, Software, and Applications (COMPSAC), 2026

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 41/100

The gist: This paper explores the use of smartphone interaction data as an objective, continuous measure of fatigue, sleep, and daily activity in patients with neurodegenerative disorders (NDD) and

Terminology

Summary

This paper explores the use of smartphone interaction data as an objective, continuous measure of fatigue, sleep, and daily activity in patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID). The authors note that current assessment of these symptoms relies on Patient Reported Outcomes (PROs) based on standardized questionnaires completed during clinical visits every few months, which suffer from subjectivity and low sensitivity to change, failing to capture day-to-day variability. The study is part of the IDEA-FAST project, which aims to identify digital endpoints for reliable assessment of NDD and IMID patients.

Data collection: Data were collected from 137 participants across 6 disease groups—Parkinson’s Disease (PD = 18), Huntington’s Disease (HD = 9), Rheumatoid Arthritis (RA = 17), Systematic Lupus Erythematosus (SLE = 17), Primary Sjogren’s Syndrome (PSS = 17), Inflammatory Bowel Disease (IBD = 17)—plus a healthy control group (HC = 42). Participants were assigned to four clinical sites (Kiel 26.95%, Newcastle 40.43%, EMC Rotterdam 19.85%, GHI Muenster 12.77%). The cohort comprised 35.46% men and 64.54% women, with a mean age near 52 years (9.2% under 30, 20.56% between 30-40, 12.05% between 40-50, 21.98% between 50-60, 19.14% between 60-70, 17.02% over 70). Ten participants used their private phones; the rest used study-provided phones. A mobile application recorded events including PHONE APP IN FOREGROUND (with app categories grouped into 11 types), PHONE SCREEN (on/off events), PHONE BATTERY, PHONE ACTIVITY, and QUESTIONNAIRE records. PROs were prompted 4 times daily (9:00, 13:00, 17:00, 21:00), each questionnaire available for 3 hours (2.5 hours for evening). PROs included physical fatigue (Feel Q1), mental fatigue (Feel Q2), anxiousness (Feel Q3), depression (Feel Q4), pain (Feel Q5), bed time, wake time, sleep quality, time to fall asleep, time awake during night, sleepiness, and physical/mental daily activities. Most questions used a 7-level Likert scale (0-6); bed/wake times used clock responses; sleep questions used drop-down menus; one activity question accepted free text. Traditional PRO surveys (FACIT-F fatigue scale and MOS-SS acute sleep scale) were completed weekly during clinical visits.

Data cleaning and preprocessing: From 150 files (137 participants), 9 files had no data and 11 files had data but no questionnaire responses—these 20 were discarded. Of the remaining 130 files, most had 10-35 days recorded (mean 23.4 days, median 26). Three files with fewer than 3 days were discarded. Gaps in data (from hours to a week) were tolerated as long as sufficient data existed. After timestamp validation, 127 files from 118 participants remained. PHONE BATTERY and PHONE ACTIVITY records were discarded; PHONE ACTIVITY was deemed too high-level compared to other project sensors. The analysis focused on screen time and app usage.

Feature extraction: Data were aggregated into 24-hour (daily) windows. App usage features were counts of events per app category (11 features). Screen events were defined as intervals between consecutive SCREEN ON and SCREEN OFF records; unmatched records were removed, and events longer than 3 hours were discarded as outliers (caused by application breaks/gaps). Screen features included mean, median, standard deviation, total screen time, number of events, maximum and minimum screen event durations. PROs were computed as daily averages of responses. Shorter windows (8-12 hours) reduced event counts and variability; longer windows (48 hours) showed no significant differences.

App distribution: Roughly 70% of applications belonged to Unknown or Utility categories. Communication (13.06%) and Wellness (12.35%) were next most common. Other categories were nearly absent. Only Unknown and Utility were well distributed across participants and windows; other categories showed irregular presence, restricting the app-usage association analysis to just these two categories.

Questionnaire coverage: Coverage (percentage of days with valid responses) varied widely: Feel Q1 (physical fatigue) 0.868, Feel Q2 (mental fatigue) 0.854, Feel Q3 (anxiousness) 0.440, Feel Q4 (depression) 0.369, Feel Q5 (pain) 0.604, To Bed Time 0.741, Woke Up Time 0.751, Sleep Details Q1 0.622, Sleep Details Q2 0.763, Sleep Details Q3 0.763, Sleepiness 0.923, Activities Q1 0.381, Activities Q2 0.382. The low coverage for Feel Q3 and Feel Q4 was surprising since they were prompted with Feel Q1 and Feel Q2 in all daily questionnaires. Activities questions (only in evening questionnaire) had coverage below 40%, suggesting the evening questionnaire was least answered.

Association analysis: The authors used repeated measures correlation (rm) to account for multiple data points per subject, which violates independence assumptions of Pearson/Spearman correlations. For app usage, only Unknown and Utility categories could be analyzed. No significant correlation was found between any app feature-PRO pair. The strongest rm was between Unknown category and Activities Q2 (rm = 0.255, ρ = 0.001), but most pairs had large significance values. For screen events, the strongest value was Total screen time-Feel Q2 (rm = 0.11, p < 0.01), which was not large enough to denote clear linear correlation. The feature-PRO pair with fewest data points was Maximum screen time event-Feel Q4 (844 points); Minimum screen time event-Sleepiness had the most (2124 points).

Cohort-based analysis: Repeating the analysis by cohort group revealed much stronger associations, confirming that whole-corpus analysis masked cohort-specific patterns. Top associations included: # events − To Bed Time in HD (rm = −0.444, ρ < 0.01, 61 points, 8 participants), Screen time std − Feel Q2 in HD (rm = 0.416, ρ < 0.01, 76 points, 8 participants), Total screen time − Activities Q1 in RA (rm = 0.354, ρ = 0.05, 44 points, 15 participants), Mean screen time − Sleep Details Q2 in HD (rm = 0.331, ρ = 0.014, 64 points, 8 participants), Maximum screen time − Feel Q2 in HD (rm = 0.327, ρ < 0.01, 75 points, 8 participants), Median screen time − Activities Q2 in SLE (rm = −0.306, ρ = 0.02, 70 points, 13 participants), Total screen time − Feel Q4 in PSS (rm = 0.305, ρ < 0.01, 130 points, 15 participants), Mean screen time − Feel Q3 in RA (rm = −0.304, ρ < 0.01, 76 points, 15 participants), # events − Activities Q2 in PD (rm = 0.3, ρ = 0.02, 66 points, 16 participants). Four associations with rm > 0.3 were in the Huntington Disease group, including the strongest overall. However, none exceeded 0.5, so no clear linear correlation could be established. The authors caution that low data points per cohort may affect results.

Demographic analysis: No significant differences were found between genders. Slightly higher correlations appeared in some age groups: Median screen time − Activities Q2 (rm = 0.27, p < 0.01) in the 20-30 age group, and Total screen time − Activities Q1 (rm = 0.29, p < 0.01) in the 60-70 age group. Other age groups showed no such elevated values.

Personal phone analysis: For the 10 participants using their private phones (9 from RA and PD cohorts), the strongest associations were: Median screen time − Feel Q1 (rm = 0.2302, ρ < 0.01), Median screen time − Feel Q2 (rm = 0.1685, ρ = 0.03), Total screen time − Feel Q1 (rm = 0.1684, ρ = 0.04), Maximum screen time − Feel Q5 (rm = −0.1679, ρ = 0.09), Minimum screen time − Feel Q5 (rm = −0.1587, ρ = 0.09), Median screen time − Woke Up Time (rm = −0.150, ρ = 0.05). These values were higher than the whole-corpus maximum (rm = 0.11) but lower than cohort-based values. The authors hypothesized that using study phones may affect interaction patterns, but the small sample size limits conclusions.

Conclusions: Low correlation values were found between mobile data and PROs when analyzing the whole corpus. Cohort-based analysis yielded higher correlations, but none strong enough to confirm a clear linear relationship. The authors suggest further analysis using more complex mobile features such as touchscreen dynamics or data from other smartphone sensors, which have shown rich information about smartphone interaction in prior work. They also note that most patients used study phones, which may affect interaction patterns.

Improvements for AI systems

Based on the scientific paper, here are specific improvements I can make to AI systems, along with what the improved system can do:


  • Improvement: Replace standard Pearson/Spearman correlation with repeated measures correlation (as described in Section II-B and III-C) in AI models that analyze paired longitudinal data from multiple subjects.

  • What the improved AI system can do: Correctly handle non-independent data points (e.g., multiple daily measurements from the same patient) without violating statistical assumptions, avoiding misleading aggregation. This is critical for digital phenotyping, wearable health monitoring, and clinical trial analysis where each subject contributes many time-points.

  • Improvement: Incorporate cohort-group-specific correlation patterns (Section III-D) into AI models, rather than using a single global model. For example, the paper found stronger associations in Huntington’s Disease (e.g., number of screen events vs. bed time, r m = -0.444) that were masked in the full dataset.

  • What the improved AI system can do: Automatically stratify users by disease group, age, or device type (personal vs. study-provided) and train separate models or apply cohort-specific weighting. This enables more accurate fatigue/sleep/activity prediction for rare or heterogeneous patient populations, avoiding the dilution of signals that occurs with pooled analysis.

  • Improvement: Implement the cleaning rules from Section III-A: discard files with 3 hours as outliers, and validate timestamp consistency.

  • What the improved AI system can do: Automatically preprocess raw smartphone logs to remove noise and artifacts (e.g., app crashes, gaps, long screen-on events) before feature extraction. This reduces false correlations and improves the reliability of downstream predictive models for mHealth applications.

  • Improvement: Use 24-hour aggregation windows and daily averaging of PROs (Section III-B) to handle sparse and irregular mobile event data, and to compensate for missing questionnaire responses.

  • What the improved AI system can do: Generate robust daily-level features (e.g., total screen time, number of events, median event duration) even when raw event frequency is low. This allows AI models to work with real-world, incomplete data from patients who do not respond to every prompt, improving coverage and reducing bias.

  • Improvement: Detect and exclude app categories with irregular or near-zero presence (e.g., only “Unknown” and “Utility” were usable; categories like “Communication” failed correlation computation due to sparse distribution). Implement a minimum-support threshold per category.

  • What the improved AI system can do: Avoid numerical instability and false negatives by automatically selecting only informative app-usage features. This makes the system more robust when analyzing app-usage patterns across diverse user populations.

  • Improvement: Distinguish between study-provided phones and personal phones (Section III-D). The paper found that participants using personal phones showed higher correlations (e.g., median screen time vs. physical fatigue, r m = 0.23), likely because interaction patterns are more natural.

  • What the improved AI system can do: Adjust feature scaling or model weights based on device ownership. This improves ecological validity and predictive accuracy for real-world deployment, where users are more likely to use their own devices.

  • Improvement: Incorporate age-group-specific correlation patterns (e.g., 20–30 age group: median screen time vs. mental activity, r m = 0.27; 60–70 age group: total screen time vs. physical activity, r m = 0.29).

  • What the improved AI system can do: Provide age-adaptive digital biomarkers. For example, an AI system can weight screen-time features differently for younger vs. older users when predicting fatigue or activity, leading to more personalized and clinically meaningful assessments.

  • Improvement: Automatically discard low-information data types (e.g., battery level, high-level activity states) and focus on screen events and app usage, as done in Section II-B.

  • What the improved AI system can do: Reduce computational overhead and model complexity while maintaining predictive power. This is especially useful for on-device AI in resource-constrained mobile health applications.

  • Improvement: Report both r m and rho (p-value) for every feature-PRO pair, and only retain associations with rho < 0.05 for downstream modeling (as in Tables III–V).

  • What the improved AI system can do: Avoid overfitting to noise by filtering out non-significant correlations. This ensures that only statistically validated digital markers are used for clinical decision support, reducing false alarms in patient monitoring.

  • Improvement: Since linear correlations were weak (max r m = 0.11 globally), the AI system should not rely solely on linear models. Instead, use non-linear models (e.g., gradient boosting, transformers) that can capture complex, non-linear interactions between mobile features and PROs, while still using RMC for feature selection.

  • What the improved AI system can do: Achieve higher predictive accuracy for fatigue, sleep, and activity from smartphone data, even when linear associations are weak. This is essential for real-world clinical utility where relationships are often non-linear and context-dependent.

Summary of improved system capabilities:

The improved AI system can process raw smartphone interaction logs, clean and aggregate them into daily features, stratify by cohort/age/device type, select only statistically valid and informative features, and build personalized non-linear models to objectively assess fatigue, sleep quality, and daily activity in patients with neurodegenerative and immune-mediated diseases—providing a reliable, continuous, and bias-reduced alternative to traditional questionnaires.

Abstract

Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID). The current assessment of such symptoms is usually conducted using patient reported outcomes (PROs) based on standardized questionnaires that patients usually complete every few months. This assessment protocol has raised some concerns, due to its propensity to exhibit biases derived from its subjectivity nature, or the low sensibility to changes, which may lead to a failure when trying to capture variability over time. In this work, we explore the use of smartphone data, which can serve as a proxy for how patients interact with their devices, to provide an effective, reliable, and objective assessment of the symptoms mentioned above. Our study comprises data from 137 participants belonging to 6 different disease groups, plus a healthy control group. We conducted statistical analysis based on repeated measures correlation, in which we analyze the correlation between screen-time and app-usage features with scores obtained from the PROs collected from the participants using a smartphone application.

Related papers