Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations

arXiv:2608.07385 · cs.LG, cs.AI, eess.AS, eess.SP, stat.ML · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations".

Jane: The paper was written by Ioannis N. Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker et al. from Khalifa University and University of Toronto and MIT Media Lab and Aristotle University of Thessaloniki.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Jane: Today's paper comes from a collaboration across four institutions — Khalifa University, the University of Toronto, MIT Media Lab, and Aristotle University of Thessaloniki. The authors are Ziogas, Khazaei, Taha, Al Shehhi, Khandoker, Hadjileontiadis, and Hatzinakos. It's a serious group, and the title they chose is a mouthful.

Tom: Every word in that title carries weight, honestly. Omni-modal, decomposition autoencoders, full-stack, disentangled — put together, it promises that one model can do everything a wearable sensing system needs.

Jane: "Omni-modal" is the word I want to sit with. We usually say "multi-modal" when there are two or three sensor types involved. These authors mean something bigger — arbitrarily many data streams, unified in a single model.

Lu: They actually test that ambition with thirty channels of data at once. That's inertial sensors, physiological signals, and audio — six audio components, three physiological streams, and twenty-one inertial channels. It's the kind of information flood that a modern smartwatch and phone can produce together.

Meng: That flood connects to something I see in the author list. The disciplinary mix is striking — biomedical engineering researchers like Khandoker and Hadjileontiadis working directly with signal processing people like Hatzinakos. The funding comes from Khalifa University's Healthcare Engineering Innovation Group, which fits that clinical angle.

Tom: Exactly. "Full-stack" is their way of saying the model should handle classification, fusion, disentangled representation learning, and generation in one place, instead of four separate systems each trained for a single job.

Jane: Those goals usually pull against each other. A model trained to classify activities tends to produce representations that are hard to interpret or generate from, and generative models often don't classify well. So the paper is making a strong claim.

Lalam: It is a strong claim, and it's worth reading precisely because of that. Getting one architecture to hold all of those abilities without collapsing into a jumble is a genuine open problem in representation learning. And whether they succeed matters beyond this one architecture, because wearable sensing is exactly where the data complexity is growing fastest.

Jane: Then that structure is exactly what we should look at next.

Summary: Tom: We've sized up the team and the ambition behind this paper, so let's get concrete about the architecture. On the surface it's simple: one shared convolutional encoder, one shared decoder, and no modality-specific branches. The encoder is a seven-layer convolutional network modeled on Wav2Vec2, while the decoder is a four-layer fully connected network.

Jane: Everything flows through the same pipeline. Each input becomes a time-frequency representation, and they build an anchor signal by literally summing all the modality channels together before the transform. For mixed acoustic and inertial signals, that anchor becomes a Mel spectrogram — the paper describes it as a sonified representation.

Lu: The thirty channels break down into six audio components from a filter decomposition, three physiological signals from a smartwatch, and twenty-one channels from seven tri-axial inertial sensors. Then the self-supervised loss takes over, and it does the real work. One term pulls each modality's latent representation toward the anchor, enforcing that all sensors describe the same underlying event, while a second term pushes modalities apart so they don't become redundant copies.

Meng: So alignment and separation happen at the same time. Each modality keeps its own subspace in the latent space, but everything is anchored to a shared context. That's how activity and identity get structured as separate factors, rather than tangled together.

Tom: On the HARWE dataset — thirty-five participants, nine daily activities, all thirty channels — the best variant reaches 84 point 56 percent accuracy for activity recognition and 88 point 97 percent for identity recognition in the subject-dependent setting. Those numbers beat both the transformer-based fusion baselines and the VAE-based alternatives.

Jane: For me the generative results are even more striking. The reconstruction mean absolute error improves by more than 76 percent relative to a multi-modal VAE baseline, and the distributional distance between real and synthetic data improves by about 14 percent. So the same latent space that recognizes activities can also synthesize believable sensor signals.

Lu: That's the full-stack claim made concrete. And the model stays at 4 point 1 million parameters whether you give it three channels or thirty, because the architecture itself doesn't grow with the modality count.

Meng: The subject-independent setting deserves a mention too, because there the results are more mixed. The multi-branch VAE baseline reaches 70 point 91 percent for activity recognition, while OmniDecVAE gets 64 point 56 percent. That suggests shared encoders may generalize a bit less to people never seen in training.

Jane: Still, the overall package is unusual — classification, generation, disentanglement, and a constant memory footprint in a single model. That's a real step beyond the typical task-specific wearable system.

Tom: Which raises the natural next question: what does that combination of abilities actually buy you in practice?

Improvements: Jane: We've seen what the model achieves on benchmarks, so let's think about the practical improvements it suggests. The first is architectural: because the encoder is shared and the modality structure lives in the loss, the parameter count stays flat at 4 point 1 million no matter how many sensors arrive. There are no per-modality branches to multiply.

Tom: That flat footprint is genuinely rare. Their multi-modal VAE comparison grows to 88 million parameters once you add all thirty channels, since every modality gets a dedicated branch. OmniDecVAEs avoid that scaling problem entirely, and that changes what you can deploy.

Lu: The paper puts real numbers on deployment too — about four milliseconds of inference latency per sample and roughly 15 megabytes on disk. That fits comfortably inside the real-time window for edge devices like smartwatches and clinical monitors, which is where wearable eye has to live.

Meng: On the generative side, the improvements open up applications that classification-only systems can't touch. You could reconstruct a sensor that failed mid-recording, synthesize extra training data for other models, or reduce the number of physical sensors by generating the modalities you stopped collecting. All of that comes from the same learned representation.

Jane: The privacy angle is the one that pulls me in. Because identity is disentangled into its own subspace, you could anonymize a recording by manipulating that subspace directly while preserving the activity content. In patient-centric healthcare, you could also keep identity available when clinicians genuinely need it — and block it when they don't.

Tom: There's a surprising video result as well. OmniDecVAE never sees video, yet in the subject-independent scenario it reaches 64 point 56 percent activity accuracy, close to the supervised transformer baseline that does use video and gets 68 point 5 percent. Video is normally the most informative modality in human activity recognition, so that gap is remarkable.

Lalam: The paper is also honest about what still needs improving — generating raw time-domain signals instead of time-frequency representations, because that would give downstream processing more freedom, plus handling missing modalities explicitly and moving beyond time series into video and text. Those are natural next steps, and the authors name them directly. That honesty makes the current claims easier to trust.

Jane: So the improvements aren't just incremental accuracy gains. The paper is proposing a different shape for wearable eye, where a single model carries recognition, generation, and privacy controls together. That's a bigger statement than any single benchmark result.

Tom: That framing goes straight back to the opening pages, where the problem is first laid out. Let's look at that framing next.

First page: Tom: We've covered what the model does and what it enables, so let's go back to the opening pages, because the framing there matters. The abstract makes a pointed claim: no existing approach operates as a full-stack wearable processor. None of them simultaneously handle classification, disentangled representation learning, fusion, and generation.

Jane: That's the gap the paper is filling. And it sits on top of a generative story — a latent sensing event, decomposed into modality-specific latent variables, which in turn produce the sensor measurements you observe. That three-step process is the conceptual backbone.

Lu: The clever part is that the data arrives pre-decomposed. You already have separate sensor streams, so instead of learning to decompose a single signal, the model learns to compose — reconstructing the underlying latent event from those separate views. That inversion is what makes the shared encoder viable.

Meng: Composition is a nice way to put it. And that decomposition loss comes with a weighting scheme which deserves more attention than it usually gets. The authors don't treat all sensor pairs equally — audio, being complex and high-dimensional, gets a stronger pull toward the anchor, while closely related channels like BVP and EDA get weaker repulsion in the orthogonality term.

Jane: Those weighting factors — 0 point 25 for the positive interactions and 0 point 4 for the negative ones — encode domain knowledge about which sensors should behave similarly. That's how the disentangled structure ends up matching the physical reality of the wearable setup. It's structure with a purpose, not structure for its own sake.

Tom: The contributions listed on that page are worth holding onto. Scalable fusion without transformer-style architectures, multi-modal time-frequency generation through a shared decoder, and a latent space where modality, activity, and identity are clearly structured — with the downstream results to back it up. The paper reports accuracy improvements of 1 point 01 percent in activity recognition and 6 point 75 percent in identity recognition over the comparison methods.

Lu: The clinical motivation is right there on page one as well. Hospitals upgrade equipment constantly, and a modality-invariant model can adapt to sensor replacements without full retraining. Combined with identity isolation for patient privacy and generation for sensor reduction, that becomes a sustainability story for healthcare eye.

Lalam: So the first page is really the blueprint for the whole paper — problem statement, generative assumption, method outline, contributions, and the healthcare motivation tying it together. Everything in the experiments follows from that opening framing. That's good scientific writing, honestly.

Jane: Which brings us to the closing question: what does this all amount to as a research direction?

Conclusion: Tom: Let's pull it all together. This paper makes a strong case that wearable eye doesn't need separate systems for recognition, generation, fusion, and privacy. One model with a shared encoder and a carefully structured loss can carry all of those, and the evidence is in the tables.

Jane: The evidence is concrete. The best variant reaches 84 point 56 percent activity accuracy and 88 point 97 percent identity accuracy in the subject-dependent setting. Reconstruction mean absolute error improves by 76 point 84 percent, distributional similarity improves by 13 point 85 percent, and the whole model runs at 4 point 1 million parameters with about four milliseconds per sample.

Lu: We should keep the caveat we discussed earlier in view. In the subject-independent setting, the multi-branch baseline edges it out on activity recognition — 70 point 91 percent versus 64 point 56 percent. So the shared-encoder design has a trade-off, and the authors acknowledge it directly.

Meng: On the other side of that caveat, they also lay out clear next steps. Raw time-domain generation, cross-modal inference with missing sensors, and expanding beyond time series are all flagged as future work. Those don't take away from the current results — they show where the framework needs to go.

Lalam: The bigger picture, from where I sit, is that structure can live in the learning objective instead of the architecture. You keep the network simple and push the domain knowledge into the loss, and that's an idea that transfers well beyond wearables. Representation learning papers often point this direction, but few demonstrate it this cleanly on real physiological and inertial data.

Jane: And the practical implications stack up — biometric security, anonymization through the identity subspace, sensor reduction, synthetic signal generation for healthcare, all in a lightweight edge-ready package. Those are exactly the problems next-generation wearable systems need to solve.

Tom: That's a lot of value from one framework, and a good note to end on. Thanks to the authors, and to everyone listening. We're ready to take the next paper from the pile.

Jane: Until next time — keep listening, and keep asking what your models are actually learning. That question is what drives work like this.

Ioannis Ziogas, Ensieh Khazaei, Bilal Taha, Aamna Al Shehhi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis, Dimitrios Hatzinakos

Khalifa University · University of Toronto · MIT Media Lab · Aristotle University of Thessaloniki

cs.LG, cs.AI, eess.AS, eess.SP, stat.ML

Submitted: 2026-08-07

Updated: 2026-08-10

Comments: 15 pages, 7 figures, 7 tables

Code: https://github.com/GiannisZgs/OmniDecVAEs

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

Key concepts

Omni-modal
Refers to handling arbitrarily many data streams in one model, unlike multi-modal which typically handles two or three. The paper tests this with 30 channels including inertial, physiological, and audio signals, all processed through a single shared pipeline.
Disentangled representation learning
A method where the model separates different factors (like activity and identity) into distinct subspaces in the latent space. This allows manipulating one factor (e.g., anonymizing identity) without affecting others, and is achieved through a loss that aligns modalities to a shared anchor while pushing them apart.
Full-stack wearable processor
A model that performs multiple tasks—classification, fusion, disentangled representation learning, and generation—in one system, instead of separate models for each. The paper claims OmniDecVAE achieves this, with a constant 4.1 million parameters regardless of sensor count.

Terminology

Summary

Published: arXiv:2608.07385v1 [cs.LG], 7 Aug 2026, IEEE Transactions and Journals Template

The paper introduces Omni-modal Variational Decomposition Autoencoders (OmniDecVAEs), a framework for learning multi-purpose representations from arbitrarily many modalities. The authors state: "Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation learning, fusion, and generative modeling of highly heterogeneous multi-modal time series. OmniDecVAEs extend DecVAEs by learning modality-conditioned time-frequency latent subspaces through a multi-view self-supervised decomposition loss and a shared asymmetric autoencoder (AE) architecture. On a challenging omni-modal human activity recognition (HAR) setting with up to thirty modalities, the method demonstrates full-stack representation ability: OmniDecVAEs full-stack disentangled representation properties lead to accuracy improvements of 1.01% and 6.75% in activity and identity recognition, respectively, compared to transformer-based and VAE-based methods. Additionally, OmniDecVAEs synthesize realistic omni-modal time-frequency data that manifest with enhanced reconstructions (mean absolute error improves by 76.84%) and distributional similarity between real and synthetic data (maximum mean discrepancy improves by 13.85%). The framework is described as a lightweight model suitable for intelligent edge wearables and clinical healthcare, unifying processing requirements and abilities in a single model, through its enhanced representational capacity, modality-invariant spatial complexity (4.1M parameters), and real-time latency."

The authors motivate the work by noting that "Wearable sensing systems are increasingly generating large-scale, heterogeneous multi-modal data, requiring robust and expressive representation learning for applications in healthcare, human-computer interaction, and biometric security. They describe a transition from conventional multi-modality to omni-modality, defined as the unification of arbitrarily many heterogeneous data streams into universal and expressive representations through a single model."

They argue that existing supervised multi-modal approaches require large amounts of labeled data, often leveraging transformer or state-space models to capture long-range dependencies, but fusing heterogeneous streams such as inertial, physiological, and behavioral data requires significant per-modality design considerations. Self-supervised learning (SSL) has emerged as an alternative, but existing approaches primarily focus on task-specific performance and do not jointly address disentanglement and generative modeling. The key gap: To the best of our knowledge, no existing approach jointly addresses scalable multi-modal fusion, disentanglement, and generation within a single unified framework.

The OmniDecVAE model processes up to thirty modalities in our experiments—including inertial, physiological, and audio signals, all transformed into the time-frequency (TF) domain using the short-time Fourier transform (STFT) and Mel-scale representations for audio. A "shared convolutional encoder learns modality-specific latent subspaces, while a shared fully-connected (FC) decoder establishes sample-level correspondence and reconstructs multi-modal signals. The variational formulation enforces a Gaussian latent structure, enabling stochastic generation of unseen samples. The model is evaluated on the HARWE dataset, outperforming supervised transformer-based fusion methods, VAE-based models, and classical approaches such as Independent Component Analysis (ICA) and Principal Component Analysis (PCA)."

The contributions are: (1) a novel multi-modal SSL objective that enables scalable fusion of arbitrarily many modalities within a structured latent space, using a modality-agnostic architecture; (2) a unified architecture that scales to a large number of modalities through a shared encoder, without relying on transformer-based designs, and enables multi-modal TF data generation through a shared decoder; and (3) a disentangled latent representation where modality, activity, and identity factors are well structured, leading to markedly improved performance in downstream recognition tasks.

The authors review multi-modal fusion in wearable computing, noting that existing fusion methods do not naturally scale to an arbitrary number of modalities without architectural modifications, limiting their applicability in omni-modal wearable settings. In SSL, they observe that in most existing methods, SSL objectives primarily enhance intra-modal representations, while cross-modal fusion is handled through architectural design rather than the learning objective itself. For disentangled representation learning and generation, they note that previous AE-based and GAN-based methods achieve high-quality signal synthesis but often do not evaluate the utility of the learned latent representations for downstream tasks.

The paper defines a multivariate time-series dataset X = x1,..., xN ∈ R T×M, where each observation is generated by a three-step generative process θ, involving a global latent variable z and a set of modality-specific latent variables z1, z2,..., zM that act as components of z. The latent variable z is interpreted as a sensing event, while each zm represents a modality-specific view of that event. The latent components are assumed orthogonal: ⟨zi, zj⟩ = Σ zi zj* = 0 for all i ≠ j, and z = Σ zi. In the multi-modal setting, the dataset is already decomposed into its constituent modalities, so instead of performing decomposition, the objective becomes to learn a composition model that reconstructs how the latent event z is expressed in the input space.

The framework operates as follows: (1) the anchor x0 = Σ xi is constructed; (2) multi-modal observations and the anchor are fed to a shared encoder fϕ to obtain latent representations hi and h0; (3) intermediate representations are mapped to modality-specific latent variables zi ∼ qϕ(zi x, z); (4) the self-supervised decomposition objective enforces alignment of hi with h0 via Llat recon and orthogonality via Lortho using omni-modal weighting; (5) modality-specific subspaces are combined into z = [z0, z1, z2,..., zM]; (6) latent variables are sampled via reparameterization and reconstruct inputs using the conditional decoder pθ(x z, c); (7) the model is trained using the conditional DELBO objective.

The SSLDec loss extends DecVAEs to the multi-modal setting. The latent reconstruction term is:

Llat recon = Σi wi pos D KL(D JS(hi, h0) ∥ p)

which minimizes the divergence between each modality-specific representation hi and an anchor representation h0, encouraging alignment across modalities. The orthogonality term is:

Lortho = Σi Σⱼ w(i,j) neg D KL(D JS(hi, hⱼ) ∥ n)

which maximizes the divergence between all pairs of modality representations, promoting separation and reducing redundancy. The use of KLD formulates this adversarial objective into stable cross-entropy terms and mitigates representation collapse. The anchor is defined in input space as x0 = Σi xi.

To handle heterogeneous relationships between channels and modalities, the authors extend SSLDec with an asymmetric contrastive formulation using omni-modal vector VOM pos and matrix WOM neg, with imbalance factors IFpos = 0.25 and IFneg = 0.4. Since the anchor representation is constructed as the sum of all input signals, modalities with higher signal complexity or dimensionality, such as audio, tend to exert a stronger influence on the anchor, so stronger weights are assigned to such anchor–modality interactions. For negative interactions, "smaller weights are assigned to pairs of channels that are expected to be similar, such as physiological signals (e.g., BVP, EDA, temperature), corresponding axes of inertial sensors, and channels originating from the same sensor, while dissimilar modalities are assigned stronger repulsive weights. This formulation can scale to an arbitrarily large number of channels and modalities... without requiring modifications to the underlying model architecture."

The encoder-only objective is:

LDELBOEnc = Llat recon − Lortho − βLprior + const.

The encoder-decoder objective is:

LDELBOEnc−Dec = E[log pθ(xnzn)] − Llat recon − Lortho − βLprior

where β follows the β-VAE formulation.

The architecture uses a seven-layer one-dimensional CNN based on Wav2Vec2, equipped with LayerNorm normalization and GELU activations, followed by a shared FC projection head mapping into a latent interaction space where SSLDec is applied. Modality-specific latent subspaces Zi are obtained via M + 1 FC projection layers that output the mean µ and variance logvar parameters of Gaussian distributions. An aggregation function concatenates subspaces into the final disentangled representation. "A key property of OmniDecVAE is that multi-modality is handled entirely through the SSLDec objective, allowing the network architecture to remain simple and modality-agnostic. No modality-specific encoder branches are required."

The decoder is a shared four-layer FC network with LayerNorm normalization and GELU activations, with no non-linearity at the output layer. The anchor latent z0 is not provided to the decoder. Reparameterization uses a scaling factor τ in the range [0.1, 1.0], where smaller values of τ... lead to more stable training as the number of modalities increases. A conditional embedding c encodes the channel–modality combination, and the decoder is optimized with a smoothed mean absolute error (sMAE) loss.

The HARWE dataset consists of video, audio, inertial and physiological recordings from thirty-five participants performing nine different daily activities in work environments. The video modality is not used. The modalities include: audio sampled at 16kHz decomposed via Filter Decomposition into six frequency band-delimited OCs; physiology including BVP (64Hz), EDA (4Hz), and Temperature (64Hz); and seven inertial measurements (ACC, GYR, GRA, OR, RO, MG, LACC), each tri-axial, totaling 21 inertial channels. The total number of modality channels is 30. Two partitioning schemes are used: Easy (subject-dependent, 80/20 split) and Difficult (subject-independent, 70/30 split).

All physiological and inertial modalities are transformed via a 512-point STFT in the frequency domain with a 0.75 s Hamming window, with an overlap of 75%, resulting in a spectrogram of size 64 × 3 that is flattened to a final size of 192 × 1. Audio uses a Mel spectrogram with 64 Mel scales and 3 hops. The anchor signal is computed by superposing all modalities in the time domain before taking a TF-domain transform... we take the Mel transform for the anchor, resulting in a cumulative 'sonified' Mel representation. Models are trained on a single NVIDIA RTX 6000 Ada with PyTorch, using Adam, for approximately 120 epochs. The learning scheme consists of pre-training with one of Eqs. (14), (15), or (17), followed by post-training with a simple SVM classifier as a classification oracle.

In the subject-dependent scenario, the β-OmniDecVAEEnc−Dec variant provides the best performance in terms of HAR and IR, achieving 84.56% HAR accuracy and 88.97% IR accuracy, closely followed by VAE-based methods. Supervised transformer-based methods have subpar performance and manifest with a much higher sensitivity to the random seed. In the subject-independent scenario, the β-MMVAE variant outperforms other alternatives, signifying that the multi-branch architecture is slightly more adaptable to unseen subjects compared to the proposed OmniDecVae.

TSNE visualizations (Fig. 3) show that "β-OmniDecVAE achieves very clear separation between different modality modes... in contrast to β-MMVAE, attributed to the LDELBOEnc−Dec objective... this disentangled modality structured latent space, renders the representation more informative to the downstream tasks of HAR and IR."

OmniDecVAEs are characterized by enhanced reconstruction quality as evident by the MSE and MAE metrics. The β-OmniDecVAE achieves MSE of 0.116 vs. 4.358 for MMVAE, and MAE of 0.236 vs. 1.019 for MMVAE. Additionally, OmniDecVAE models manifest with a higher distributional similarity to real data. Namely, a lower MK-MMD and DivScore suggest that the distribution learned by OmniDecVAE is more 'realistic', closer to real data (MK-MMD 0.181 vs. 0.210; DivScore 4.094 vs. 15.468). Figure 4 visually confirms that OmniDecVAEs provide much more accurate reconstructions w.r.t. the amplitude and position of TF events, whereas they capture better the morphology of narrow-band frequency phenomena such as the decomposed audio events.

Modality ablation (Tables IV, V) shows that "the full-modal scenario with C = 30 channels does not provide the best performance in the SD scenario, but rather the combinations with less modalities. Specifically, combinations that contain the inertial sensors demonstrate superior performance... In the SI scenario though, the full-modal scenario provides the best results overall. Figure 5 shows β-OmniDecVAE demonstrates a more robust behavior, less influenced by the number of input modality channels, compared to β-MMVAE."

Loss component ablation (Table VI) reveals that "utilizing the SSLDec terms alongside a supervised loss, gives the highest performance while maintaining disentanglement between activity and identity. It becomes evident that SSLDec and the decoder are the main drivers of disentanglement."

The β ablation (Fig. 6) shows that for the full model, "higher β values result in a collapse of the disentangled structure. We also see that the absence of the Lprior through β = 0 is not detrimental to the downstream performance, mainly attributed to the disentanglement mechanisms of SSLDec. A different behavior is evident in the decoder-less variant... in the absence of a decoder, a higher β is needed for a disentangled representation."

OmniDecVAE provides the lighter alternative in terms of SoD and number of parameters (15.82MB and 4.11M parameters, versus 336.46MB and 88.09M for MMVAE), while MMVAE presents increasing storage requirements due to the high number of modalities it accommodates as separate branches. Figure 7 shows the spatial complexity of MMVAEs increase with the number of modalities; on the other hand OmniDecVAEs are invariant to the number of modalities for that matter. Although OmniDecVAE requires more FLOPs (85.85 GFLOPs) and higher latency (4.00ms) than compared methods, it still lie[s] within the real-time processing window for edge applications, with a latency of 4ms per sample.

The authors note that "even though our OmniDecVAE approach does not utilize the video modality, one of the most expressive modalities in HAR, it performs on par with a supervised transformer-based method in the HARWE Difficult SI scenario that utilizes video (Acc. of 68.5% in SI HAR). They emphasize that our omni-modal AE networks performed better under SSL pre-training on both HAR and IR, compared to supervised transformer-based alternatives... and VAE-based methods and other benchmarks. The structure-informed objective is fully explainable and transparent, and the complexity-friendly design is exemplified when the number of modalities aggressively scales; architectures with separate branches per modality quickly become unsustainable in real-life edge deployment scenarios."

Future work includes generating synthetic data in the more raw format of time domain, cross-modal inference and missing modality scenarios, and accommodating other types of modalities beyond time series, such as video or text.

The conclusion states: "OmniDecVAEs simultaneously enhance HAR and IR accuracy by 6.75%, and 1.01%, multi-modal data synthesis MAE and MK-MMD by 13.85% and 76.84%, respectively, along with modality-invariant storage requirements of 15MB or 4.1M parameters. These results underscore OmniDecVAEs as a foundational paradigm for next-generation full-stack models in wearable AI."

Improvements for AI systems

Based on the paper, OmniDecVAEs enable the following concrete improvements to AI systems:

1. Arbitrary-modal scalability without architectural redesign

  • The shared encoder–decoder architecture replaces per-modality branches, so AI systems can ingest new sensor streams (e.g., adding a new physiological or inertial channel) without changing the model structure.

  • Spatial complexity remains constant at 4.1M parameters / 15MB regardless of modality count, enabling deployment on edge wearables instead of cloud servers.

2. Disentangled, interpretable latent representations via self-supervised decomposition

  • A new SSL objective (SSLDec) explicitly decomposes the latent space into orthogonal modality-specific subspaces plus a shared anchor subspace.

  • The improved AI system can separate latent factors for modality, activity, and identity, making representations more informative for downstream tasks:

  • Human activity recognition accuracy increases by 1.01% (84.56% on HARWE)

  • Identity recognition accuracy increases by 6.75% (88.97%)

  • T-SNE visualizations show clear separation between modality modes, so latent structure is more explainable and transparent.

3. Stronger generative modeling of multi-modal time–frequency signals

  • The variational decoder enables stochastic generation of synthetic omni-modal data.

  • Reconstruction quality improves dramatically compared to MMVAE:

  • Mean squared error drops from 4.358 to 0.116

  • Mean absolute error improves by 76.84%

  • Distributional realism improves:

  • Maximum mean discrepancy improves by 13.85%

  • Divergence score drops from 15.468 to 4.094

  • The system can therefore synthesize realistic multi-modal sensor data for data augmentation, simulation, and privacy-preserving data sharing.

4. Omni-modal weighting for heterogeneous channel relationships

  • Asymmetric contrastive weights handle imbalance: high-complexity modalities (e.g., audio) get stronger positive alignment with the anchor; similar channels get weaker negative repulsion; dissimilar channels get stronger repulsion.

  • This allows the AI system to manage 30 heterogeneous channels (audio, BVP, EDA, temperature, 21 inertial channels) in one unified model without manual per-channel feature engineering.

5. Stable training via variational decomposition and reparameterization

  • Replacing adversarial objectives with KLD-based cross-entropy terms mitigates representation collapse.

  • A scaling factor τ in reparameterization stabilizes training as modality count increases.

  • Pre-training with SSL + post-training with a simple SVM classifier outperforms supervised transformer-based fusion, reducing dependence on large labeled datasets.

6. Real-time edge inference

  • The system achieves 4ms latency per sample and 85.85 GFLOPs, fitting within real-time processing windows for wearable healthcare and biometric security applications.

7. Robustness and generalizability

  • Modality ablation shows the model is less sensitive to the number of input channels than multi-branch VAE models.

  • Subject-independent evaluation shows it can generalize to unseen individuals, and even without video input it matches a supervised transformer method that uses video, demonstrating robustness to missing modalities.

Sources

Related papers