Representation Matters in Longitudinal Affective Computing

arXiv:2608.07518 · cs.HC, cs.AI · Submitted 2026-07-02 · Read on arXiv

Igor Matias, Maximilian Haas, Eric J. Daza, Matthias Kliegel, Katarzyna Wac

University of Geneva · UniDistance Suisse · Stats-of-1 · Boehringer Ingelheim Pharmaceuticals Inc.

cs.HC, cs.AI

Submitted: 2026-07-02

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

The gist: This paper addresses the temporal mismatch between dense, continuous wearable sensor data and sparse, episodic affective and cognitive labels in longitudinal, in-the-wild studies.

Terminology

Summary

This paper addresses the temporal mismatch between dense, continuous wearable sensor data and sparse, episodic affective and cognitive labels in longitudinal, in-the-wild studies. The authors recast this cadence mismatch as a temporal representation problem and compare three wave-level mappings from dense histories to sparse labels: levels (within-wave summaries), absolute drift (change across waves), and proportional drift. Using almost a year of data from 82 adults in the Providemus alz study, they model 21 affect and cognition outcomes. Day-scale signals are reduced to compact wave-level descriptors (central tendency, dispersion, and distributional shape) and learned with four regressors under two orthogonal evaluation axes—leave-one-subject-out and leave-one-wave-out. Performance is reported as scaled MAE using both mean and median across folds.

The paper's key findings are that affective states are best predicted by wave-to-wave absolute drift, whereas cognitive performance aligns with within-wave levels, reflecting emotion dynamic theories. Across windowing features, shape descriptors (e.g., minima, kurtosis) carry more signal than simple means/medians. Specifically, the most correlated metrics were Kurt for models using Levels, IQR for models using ΔABS across waves, and Min for models using the Δ%, while the least used metrics were the M, SD, and Mdn.

The authors contribute three main outputs: "(i) Evidence that distributional 'shape' features—kurtosis, IQR, and minima—outperform means and medians when day-long sensor traces are reduced to wave-level predictors. (ii) A representation triad that any intensive longitudinal study can adopt as a first-pass diagnostic. (iii) A dual cross-validation strategy and comparison that separates within-person interpolation from true out-of-sample generalization—essential for real-world deployment."

The paper concludes that "By showing precisely when change matters more than state—and when it does not—we move affective computing from passive observation toward proactive, personalized care, bringing the field a step closer to unobtrusive, always-on systems for monitoring affective states and cognitive health."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, and what the improved systems can do:


Improvement: Add a pre-processing layer that automatically selects between three representations—Levels (within-wave summaries), ΔABS (absolute drift), and Δ% (proportional drift)—based on the target outcome type (affective vs. cognitive).

What the improved AI system can do:

  • For affective state prediction (e.g., stress, anxiety, depression), automatically switch to ΔABS features, improving prediction accuracy by up to 20% compared to using Levels alone (as demonstrated in the paper).

  • For cognitive performance prediction (e.g., memory, processing speed), automatically use Levels, avoiding the noise introduced by drift-based features.

  • Eliminate the need for manual feature engineering decisions, reducing model brittleness and improving reproducibility across studies.

Improvement: Replace default mean/median summarization with a richer set of distributional descriptors: kurtosis, IQR, minimum, maximum, skewness—prioritized by the paper’s findings that shape features carry more signal.

Improvement: Implement a mandatory dual evaluation: LOSO (leave-one-subject-out) for cross-participant generalization and LOWO (leave-one-wave-out) for temporal robustness.

Improvement: Build a personalized baseline tracker using ΔABS features for affective outcomes, with a threshold based on the participant’s historical wave-to-wave variability.

Improvement: Implement a rolling baseline of Level-based metrics (e.g., monthly IQR of sleep efficiency, resting HR) for cognitive outcomes, with a persistent downward drift trigger.

Improvement: Replace single-pass feature selection with a two-pass approach: retain features explaining ≥90% cumulative SHAP (for tree models) or positive permutation importance (for SVM).

Improvement: Include a Δdays control variable in all ΔABS and Δ% models to adjust for self-selected assessment times.

Improvement: Implement the paper’s data quality rules: require ≥50% valid days per wave, ≥10 hours wear time per day, and treat zero-value HR as missing; no imputation.

Improvement: Use the paper’s Table III to pre-select the best representation per outcome (e.g., ΔABS for all PROs, Levels for most PerfROs) before training.

Improvement: Expose a compact feature schema (8 statistics × 38 sensors = 304 features max) that can be computed locally and transmitted as a single vector per wave.

Summary of Capabilities: The improved AI system can now predict affective states and cognitive performance with higher accuracy by automatically choosing the correct temporal representation, using shape-aware features, and validating under both cross-participant and cross-time protocols. It can operate on-device, trigger personalized just-in-time interventions, and monitor cognitive decline over months—all while maintaining privacy and reducing participant burden.

Abstract

Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are labeled only episodically (per waves). We recast this cadence mismatch as a temporal representation problem and compare three wave-level mappings from dense histories to sparse labels: levels (within-wave summaries), absolute drift (change across waves), and proportional drift. Using almost a year of data from 82 adults in the Providemus alz study, we model 21 affect and cognition outcomes. Day-scale signals are reduced to compact wave-level descriptors (central tendency, dispersion, and distributional shape) and learned with four regressors under two orthogonal evaluation axes: leave-one-subject-out and leave-one-wave-out. Performance is reported as scaled MAE using both mean and median across folds. Differences emerge: affective states are best predicted by wave-to-wave absolute drift, whereas cognitive performance aligns with within-wave levels, reflecting emotion dynamic theories. Across windowing features, shape descriptors (e.g., minima, kurtosis) carry more signal than simple means/medians. We contribute a representation triad for sparse-label modelling, a wave-level feature schema applicable on-device, and a dual-axis reporting practice that separates cross-participant generalization from temporal robustness. These results convert temporal representation from an implicit preprocessing step into an explicit, testable design choice for real-world affective-computing applications in brain health.

Related papers