Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals".
Jane: The paper was written by Wanting Mao, Maxwell A Xu, Harish Haresamudram, Mithun Saha, Santosh Kumar et al. from University of Illinois Urbana-Champaign and University of Memphis.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, we’ve seen the title, but let’s talk about what Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals actually does under the hood.
Jane: The core idea is that existing methods often rely on a contrastive approach which tends to focus only on the things both working together—the "between-modality" stuff.
Lu: And they risk losing the unique signatures of each sensor, which is what they call "within-modality information," because of that reliance on simple alignment.
Meng: That’s a huge practical problem; if we lose the unique signature, we might miss critical details, like a specific type of gait or even more detail in heart rate fluctuation.
Lalam: It’s about ensuring the AI isn't just finding what's easy to align but actually capturing the full richness of human physiology.
Tom: The authors propose ProtoMM to address this by using a shared prototype dictionary instead of just negative sampling, which is a major conceptual shift.
Jane: Instead of forcing two things to be similar if they look alike, ProtoMM uses these prototypes as anchors that allow both signals to be clustered around them.
Lu: This allows the system to learn features that are consistent across modalities while simultaneously preserving the unique characteristics of each modality.
Meng: It suggests a much more sophisticated training objective than what's currently standard in many multimodal AI architectures.
Lalam: This is incredibly hopeful because it implies we can build a model that understands human behavior and health with nuance, not just crude categorization.
Improvements: Tom: The paper shows some very impressive results, and this leads us to how ProtoMM improves upon previous methods through its specific design.
Jane: We've seen it outperform various baselines in the experiments, and that’s due to how it manages both within- and between-modality data simultaneously.
Lu: The mechanism is a Swapped Prediction Loss, which allows the model to become discrete anchors for the shared latent space, making sense of continuous noisy signals.
Meng: I think what’s important practically is that this performance isn't just theoretical; it's demonstrated across three different datasets for various tasks.
Lalam: It gives us confidence that we are building a system that works in the messy world, not just in a controlled lab environment.
Tom: The paper suggests that by balancing these two loss components, we can achieve better results than if we were to optimize for only one of them at all.
Jane: That ability to balance is key; instead of emphasizing alignment or emphasizing uniqueness, it does both.
Lu: It’s a sophisticated way of saying that the optimal solution lies in finding a sweet spot between the collective pattern and the individual contribution.
Meng: If we can hit that sweet spot consistently, we’re looking at a very stable foundation model for future applications.
Lalam: It ensures that as we integrate more complex biological data, our AI will be capable of handling it with precision and integrity.
Improvements (cont.): Tom: The experimental results are really validating the hypothesis that ProtoMM is a superior way to learn these signals.
Jane: Looking at the comparison, we see that even though multimodal models generally do better than unimodal ones, there’s a unique case where they perform worse in heart rate prediction.
Lu: That’s an interesting point, because it shows the complexity of how different tasks rely on different parts of the system.
Meng: The fact that ProtoMM can still achieve the best performance overall is encouraging; it suggests robust generalization across tasks like stress detection and activity recognition.
Lalam: This robustness is vital for creating a trustworthy AI in a field where mistakes could have real-world consequences for human health.
Tom: The paper also shows that when they test different settings, the number of prototypes doesn't need to be massive to get great performance.
Jane: That’s comforting; it means we don're not building a system that requires an enormous amount of computational resources just to function properly.
Lu: It validates the idea that a focused set of meaningful anchors is more powerful than brute-forcing a vast number of potential patterns.
Meng: Having a stable, efficient architecture is what makes this practical for deployment on consumer devices.
Lalam: This stability means the technology can scale and provide reliable insights to anyone who needs them, regardless of the hardware they use.
Conclusion: Tom: We’ve covered a lot of ground today, from the initial idea to how it works and what it achieves with Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals.
Jane: It's clear that this approach is not just a small tweak; it represents a fundamental shift in how we teach AI to understand complex physiological signals.
Lu: I think the ability to see these learned prototypes, as mentioned in the appendix, really helps us visualize the underlying structure of human behavior.
Meng: That level of transparency is crucial for me; knowing how the AI makes decisions is a massive practical advantage for safety and trust.
Lalam: The future looks very bright because we are moving toward an era where personalized health monitoring isn't just possible but incredibly insightful, thanks to the insights from Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals.
Tom: It’s a powerful combination of biology, mathematics, and AI.
Jane: We hope this sets the stage for much more advanced research in personalized health tracking.
Lu: We should be seeing these models applied to other sensors and modalities very soon as well.
Meng: I'm eager to see how this scales up to multiple synchronized signals at the consumer level.
Lalam: It promises a culture where technology truly empowers individuals by understanding their own bodies better than ever before, thanks to Prototype-based Self-Supervised Multimodal Learning for PPG and Accelerometry Signals.
Wanting Mao, Maxwell A Xu, Harish Haresamudram, Mithun Saha, Santosh Kumar, James Matthew Rehg
University of Illinois Urbana-Champaign · University of Memphis
cs.LG
Submitted: 2026-08-19
Updated: 2026-08-20
Code: https://github.com/cruiseresearchgroup/COCOA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 82/100
The gist: Abstract: "Modeling multi-modal time-series data is critical for capturing system-level dynamics...
Key concepts
- ProtoMM
- The authors propose ProtoMM to overcome limitations in standard multimodal AI architectures. Instead of using simple negative sampling, it employs a shared prototype dictionary. These prototypes act as anchors that allow both signals to be clustered while simultaneously preserving the unique characteristics of each modality.
- Within-modality information
- This refers to the unique signatures or critical details found within a single sensor's data, such as specific gait patterns or heart rate fluctuations. The paper notes that traditional methods risk losing this vital information by focusing only on alignment between modalities.
- Swapped Prediction Loss
- This is the specific mechanism used in ProtoMM. It enables the model to become discrete anchors within the shared latent space. This allows the system to effectively make sense of continuous and noisy physiological signals.
Terminology
Summary
Abstract: "Modeling multi-modal time-series data is critical for capturing system-level dynamics... existing multi-modal approaches often rely on CLIP-style contrastive objectives that overfit to easily aligned features and misclassify valid cross-modal relationships as negatives, resulting in fragmented and non-generalizable embeddings. To overcome these limitations, we propose ProtoMM, a novel SSL framework that introduces a shared prototype dictionary to anchor heterogeneous modalities in a common embedding space... By clustering representations around shared prototypes rather than explicit negative sampling, our method captures complementary information across modalities and provides a coherent “common language” for physiological signals."
Introduction:
The motivation for this work stems from the need to develop "effective models for multi-modal time series biosignal data, so that complementary sensing modalities can be leveraged to overcome the ambiguities and noise that are inherent in wearable signals collected in the field environment. While previous works focused on unimodal Foundation Models (FMs), recent efforts have addressed multimodal alignment using
CLIP-style contrastive objectives that pull temporally aligned signals together. However, a key challenge exists:
When modalities are highly complementary, such as PPG and accelerometry, there is a danger that the alignment process could emphasize between-modality features at the expense of within-modality features. This risk means that focusing solely on signal alignment
could inadvertently discard information which is unique to each modality and critical for downstream tasks such as stress detection."
ProtoMM's Hypothesis:
"The goal of ProtoMM is to test the hypothesis that alignment of complementary signal modalities can be facilitated by clustering continuous and noisy biosignals into a shared set of prototypes that capture recurring and semantically meaningful patterns. We hypothesize that these prototypes will be effective in encoding within-modality features and preserving them during the embedding process."
Methodology (ProtoMM):
ProtoMM achieves its goal through a Multimodal Prototype Prediction loss,
which is defined by enforcing consistency across all pairs of views. The framework involves the following key components:
-
Problem Setup: A multimodal time-series is defined as X t = X t,1, X t,2,, X t,M where M is the number of sensor modalities."
-
Shared Prototype Space: The model introduces
a set of P trainable prototype vectors, organized as columns in a matrix P = [p 1, p 2,, p P] in R E times P.
This space isshared across all modalities and training instances.
-
Projection and Assignment: Embeddings (E m) are projected onto these prototypes yielding similarity scores S. These scores are converted into a soft assignment target U using the Sinkhorn-Knopp algorithm, which
enforces an equipartition constraint to prevent mode collapse,
and a probability distribution V m using Softmax. -
The Multimodal Prototype Prediction Loss (LMPP): This loss is designed to ensure that
any pair of views originating from the same underlying system, regardless of modality, should be able to predict each other’s prototype assignment.
It comprises two components:
-
Within-modality loss (L within-mod): Enforces consistency between all pairs of augmentations within a single modality.
-
Between-modality loss (L between-mod): Enforces consistency between all pairs from different modalities.
The final objective, the Multimodal Prototype Prediction Loss (LMPP), is a linear combination of these two losses: LMPP = 1 over A times M 2(alpha L within-mod + (1 - alpha)L between-mod).
Experimental Design and Results:
The pre-training data utilized was the initial 10 days of data from the large-scale Mobile Open Observation of Daily Stressors (MOODS) study [30].
The models were trained using the Adam optimizer with a learning rate of 10-5 and a batch size of 256.
Analysis of Benchmark Results:
-
Superior Performance:
ProtoMM achieves the strongest results,
achieving the best overall performance in4/6 downstream tasks and outperforms all other multimodal methods in 5/6 tasks.
-
Modeling Both Features: While some baselines focused on modeling only between-modality information, ProtoMM's success is attributed to its ability to model both:
The results show a clear trend where multimodal models outperform their unimodal counterparts.
-
Prototype Effectiveness: The comparison with the prototype-free analogue, SLIP, confirmed that
ProtoMM achieves superior performance over its direct prototype-free analogue, SLIP... The performance gains are particularly pronounced on the WESAD dataset.
-
Visualizing Prototypes: Qualitative analysis showed that
The shared prototypes semantically anchor the representation, unifying observable patterns and underlying physiological context to expose latent states.
Figure 2 demonstrated thatNeighbors around each prototype show consistent labels and morphology; the prototypes capture distinct patterns.
Understanding the Multi-Modality in ProtoMM:
-
Balancing Objectives: The core advantage of ProtoMM is its ability to simultaneously capture both types of information. Table 3 showed that "performance is generally higher when alpha is not 0 or 1 (best at alpha = 0.5 and 0.67), demonstrating that successful multimodal integration requires explicitly modeling both the unique contributions of each modality and their synergistic interactions."
-
Knowledge Transfer: The results showed that
both encoders (i.e., for the accelerometer and PPG) outperform their unimodal ProtoMM Within-Mod counterparts on nearly all tasks,
indicating that the shared prototype space encourages representations to besemantically aligned with the broader physiological context across modalities.
Conclusion:
"In this paper, we presented ProtoMM, a prototype-based multimodal framework for self-supervised learning on time-series data to pre-train a pulse-motion foundation model. By leveraging a shared prototype space, it aligns embeddings from different modalities and uncovers common latent physiological states... Comprehensive experiments across three datasets and six downstream tasks demonstrate that ProtoMM consistently outperforms a broad set of strong baselines."
Improvements for AI systems
Improvements to the ProtoMM Framework
Based on the architecture and findings of ProtoMM, here are specific, high-impact improvements for a next-generation AI system:
The current implementation relies on a fixed set of P trainable prototype vectors. This is sub-optimal for highly non-stationary or complex real-world data streams.
-
Improvement: Implement Dynamic Prototype Generation (DPG) using a density estimation approach (e.g., Mixture of Experts or Density Estimation Networks) alongside the existing prototypes. Instead of relying solely on a fixed P set, the model dynamically generates new candidate prototypes when input embeddings fall outside an established threshold from the nearest k existing prototypes.
-
Mechanism: When a new embedding E m is generated, if its proximity to all P is below a dynamic threshold tau dyn, it is passed through a small sub-network (e.g, an MLP) and added to the prototype dictionary P. This ensures the model can capture novel patterns not seen during initial training.
-
Resulting System Capability: The system gains superior generalization and adaptability to
out-of-distribution
(OOD) physiological events, such as sudden severe stress or unusual movement patterns, without needing retraining.
The current Multimodal Prototype Prediction Loss (LMPP) enforces consistency between pairs of views (L within-mod and L between-mod). This assumes a uniform level of semantic importance across all pairs.
- Improvement: Introduce Hierarchical Consistency Weighting (HCW) into the alpha hyperparameter structure. Instead of a fixed alpha=0.5, use a dynamic weighting scheme where the contribution of L within-mod and L between-mod is modulated by the entropy and temporal variance of the input streams.
alpha t, m = Sigmoid(beta times Var(X t,m) + gamma times H(E t,m))
Where Var(X t,m) is the variance of the m-th modality at time t, and H(E t,m) is the entropy of the embedding. High-variance/high-entropy signals (e.g., sudden movement) automatically increase the weight given to within-modality features (alpha to 1), while stable signals prioritize cross-modal alignment (alpha to 0).
-
Mechanism: The model learns to dynamically adjust its focus based on the signal's characteristics, ensuring that critical, high-energy events are not diluted by low-energy baseline data.
-
Resulting System Capability: The system achieves superior sensitivity and robustness in real-time monitoring, automatically prioritizing information during peak activity or acute stress events.
The current use of Sinkhorn-Knopp provides a stable, soft assignment target U(a) for preventing mode collapse.
-
Improvement: Replace standard Sinkhorn-Knopp with Adaptive Entropic Normalization (AEN). Instead of using a fixed number of iterations for the Sinkhorn algorithm, dynamically adjust the iteration count based on the observed
spread
or dispersion of the input embeddings E m. -
Mechanism: When input embeddings are tightly clustered (low dispersion), fewer iterations suffice, speeding up training. When inputs are widely dispersed (high variance), more iterations are required to achieve a stable, equitable assignment target U(a). This ensures optimal convergence speed and assignment quality for all data conditions.
-
Resulting System Capability: Dramatically reduces the computational overhead of the self-supervised pre-training phase while maintaining guaranteed convergence stability across varying datasets (e.g., highly synchronized vs. noisy field data).
The current system concatenates embeddings E P + E A for downstream tasks, treating the entire feature vector as a monolithic input.
-
Improvement: Implement Decomposed Prototype Probing (DPP). Instead of using the concatenated embedding, we utilize dedicated
Prototype Probes
that map specific prototype clusters back to the original modalities. For example, a subset of prototypes P ACC is designed to capture kinematic features, and P PPG captures cardiac features. Downstream tasks are then performed on these decomposed feature subspaces. -
Mechanism: The final feature vector becomes F = [Proj P ACC(E A) Proj P PPG(E P)]. This allows the the downstream classifier to explicitly query
What is the kinematic state?
orWhat is the cardiac stress level?
-
Resulting System Capability: Provides unprecedented interpretability. The system can not only predict a state (e.g.,
Stressed
) but also provide evidence for why that prediction was made (e.g.,High kinematic variability combined with high PPG heart rate acceleration
).
What the Improved AI System Can Do
A system utilizing these improvements will be fundamentally superior to current multimodal SSL approaches:
-
Achieve Superior Reliability in Edge Environments: By incorporating DPG and AEN, the system maintains high accuracy even when processing highly noisy, low-quality sensor data (e.g., movement artifacts or poor electrode contact), ensuring robust performance where traditional contrastive methods fail.
-
Provide Causal/Explanatory Insights: Through Decomposed Prototype Probing (DPP), it transforms the AI from a
black box
predictor into an explainable diagnostic tool. It can quantify the relationship between physical exertion (accelerometry features) and physiological stress response (PPG features) in real-time, providing actionable insights to clinicians or users. -
Adapt and Evolve: The system possesses intrinsic adaptability via DPG, allowing it to continuously learn from its operational environment without the massive computational cost of full model retraining, leading to a truly
lifelong learning
AI agent for health monitoring.
Sources
- Large-scale Training of Foundation Models for Wearable Biosignals
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Adam: A Method for Stochastic Optimization
- IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text
- Scaling Wearable Foundation Models
- Representation Learning with Contrastive Predictive Coding
- PaPaGei: Open Foundation Models for Optical Physiological Signals
- Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications Across Lab and Field Settings
- Exploring Contrastive Learning in Human Activity Recognition for Healthcare
- LSM-2: Learning from Incomplete Wearable Sensor Data
- SensorLM: Learning the Language of Wearable Sensors
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks