Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness

arXiv:2608.12592 · cs.LG · Submitted 2026-08-17 · Read on arXiv

Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Konz, Tianlong Chen

University of North Carolina at Chapel Hill

cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 17 pages, 5 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: Paper: "Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness" by Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Konz, and Tianlong Chen (UNITES

Terminology

Summary

Paper: Represent, Then Generate: Multimodal-Conditioned Time-Series Generation under Irregular Missingness by Haochen Zhang, Jiaheng Guo, Yu-Chao Huang, Nicholas Konz, and Tianlong Chen (UNITES Lab, University of North Carolina at Chapel Hill). Preprint under review, arXiv:2608.12592v1, August 12, 2026.

Continuous physiological time series are foundational to modern clinical monitoring, but many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. Conditional generation offers a remedy: an absent signal can be synthesized from co-recorded signals and routine clinical variables. However, existing generators... are built around a single conditioning modality and degrade when forced to handle the heterogeneous, irregularly missing mix of time-variant signals and static covariates seen in practice.

The paper identifies two key reasons why direct multimodal conditioning fails: (1) "co-recorded time series arise from distinct physiological processes with their own dynamics and clinical meaning... so an in-painting mechanism that treats them as homogeneous context cannot represent these modality-specific semantics; (2) each modality has its own irregular sampling grid and missingness pattern, making it hard to learn representations of the time-variant conditions while performing generation."

The authors propose ReCoGen (Represent Conditions, then Generate), a two-stage framework that decouples multimodal condition representation from target generation.

Stage I: Per-Modality Masked Autoencoder Training — Stage I trains one masked autoencoder per modality, distilling each time-variant condition into a compact and missingness-tolerant token sequence. Each encoder imputes from context and distills an irregular, time-variant modality into a compact, robust token sequence. The loss is mean squared error on held-out steps only: Reconstructing unseen points forces the encoder to infer from context rather than copy its input, which keeps it robust when a modality is sparsely observed. The default mask ratio is ρ=0.3.

Stage II: Conditional Flow-Matching Generator — Stage II freezes these encoders and trains a flow-matching model that synthesizes the target from both static and time-variant conditions. The target series is mapped to a 2-D image via an invertible delay embedding, and a DiT/SiT-style vision transformer is trained with x-prediction v-loss flow matching. The conditioning path uses:

  • learnable per-modality query tokens that cross-attend over its frozen autoencoder latents for time-variant modalities

  • a dual token-plus-AdaLN route for the static conditions (in-context token plus adaptive layer-norm modulation)

  1. "We show that the current methods for conditional time-series generation perform poorly under multimodal physiological conditioning; and we trace the failure to the absence of a dedicated, missingness-aware condition representation."

  2. We propose ReCoGen, a two-stage framework that decouples condition representation from generation.

  3. We conduct comprehensive experiments across three physiological datasets, showing that ReCoGen outperforms six representative conditional generators in downstream task utility.

Datasets: (i) AI-READI (type-2-diabetes cohort): generate CGM from four wearable modalities (heart rate, calorie expenditure, physical activity, respiratory rate) plus static tabular features; (ii) MIMIC-III and (iii) MIMIC-IV: generate mean ABP from three vitals (heart rate, respiratory rate, SpO2) plus static clinical vector. All windows are 24 hours at 5-minute resolution (T=288).

Baselines: Six representative conditional generators: Diffusion-TS and ImagenTime (signal-conditioned, in-painting), TimeWeaver and WaveStitch (attribute-conditioned), VerbalTS and Bridge (text-conditioned). Everything outside that interface is held fixed: all baselines reuse our dataset builder, split, windowing and valid-window filter verbatim.

Evaluation protocol: Because no ground-truth 'clean' target exists for held-out subjects, we assess generation quality by downstream clinical utility rather than by point-wise reconstruction error. A classifier trained on real signals scores each method's generated signals against true labels (train-on-real, test-on-synthetic protocol). The downstream tasks are four-class study group classification on AI-READI and external ICD diagnoses (sepsis, heart failure) plus in-hospital mortality on MIMIC-III/-IV.

ReCoGen is the strongest generative model in every row of Table 1, attaining the best AUROC and AUPRC on all sixteen (dataset, task, metric) combinations. The margin is largest on critical-care benchmarks where baselines hover near chance. On AI-READI, our generated CGM leads clearly without the real labs and matches the strongest baseline once they are fused in.

Against the Real-Valid* reference (real signal scored under the same probe), ReCoGen reaches or exceeds it on thirteen of the sixteen rows and falls below it on three, most visibly on MIMIC-III mortality (AUROC 0.599 vs. 0.658, AUPRC 0.249 vs. 0.316).

The paper notes: Downstream utility therefore measures label-relevant information transfer, and conflates target fidelity with condition re-expression. Scores above Real-Valid* are evidence that label-relevant structure survives generation, possibly in a cleaner form than the measured waveform, not that the synthetic signal is the more faithful one.

Time-series vs. static conditioning: Removing either group lowers AUROC on every task, so both carry information that ReCoGen transfers into its generation; dropping the static conditions hurts more in most cases.

Time-series conditioning mechanism: Comparing four injection methods (cross-attention tokens, in-context tokens, AdaLN vector, in-painting), Cross-attention wins overall, most clearly on the ABP tasks, where the other injections fall well behind and in-painting is unstable, dropping toward or below chance on the hardest cases.

Static-feature encoding: "Both paths together win on every (dataset, task), by a wide margin on MIMIC-IV, so they are complementary rather than redundant; between the single-path variants, removing the AdaLN modulation is consistently more damaging."

Stage-I mask-ratio sensitivity: The generator is fairly robust: AUROC varies within a modest range with no monotone trend, and 0.3 is a consistently strong setting.

"Two limitations frame that result: downstream utility conflates target fidelity with a re-expression of the conditions, so the real-signal reference is an anchor rather than a ceiling; and encoding each modality independently leaves cross-modal dependencies to the generator. Modeling the conditions jointly, and separating fidelity from re-expression in a paired setting, are natural next steps."

"ReCoGen addresses conditional generation of physiological time series from several irregularly sampled, partially observed signals together with static clinical variables... It attains the best downstream utility on all sixteen (dataset, task, metric) settings across AI-READI, MIMIC-III and MIMIC-IV, by the widest margin on the critical-care ABP benchmarks where the six baselines hover near chance."

Improvements for AI systems

Improvements to AI systems:

  1. Two-stage decoupling of condition representation from generation. Instead of forcing a single model to jointly encode heterogeneous inputs and generate outputs, the AI system first trains per-modality masked autoencoders to produce compact, missingness-tolerant token sequences, then freezes those encoders and trains a separate flow-matching generator. This prevents representation collapse when modalities have distinct dynamics and irregular sampling grids.

  2. Per-modality tokenization with masked reconstruction. Each time-varying condition is distilled into a token sequence by reconstructing only held-out time steps (mask ratio 0.3), forcing the encoder to infer from context rather than copy input. This makes the system robust to sparse or missing observations without explicit imputation.

  3. Cross-attention over frozen modality latents. The generator uses learnable per-modality query tokens that cross-attend over each modality's frozen autoencoder latents, rather than concatenating raw signals or in-painting. This preserves modality-specific semantics and avoids the instability of in-painting approaches on hard cases.

  4. Dual-path static conditioning. Static clinical variables are injected both as in-context tokens and via adaptive layer-norm (AdaLN) modulation. The two paths are complementary—removing either degrades performance, but removing AdaLN is more damaging—so the system should use both.

  5. Invertible delay embedding for target series. The target time series is mapped to a 2-D image via an invertible delay embedding, allowing a vision transformer (DiT/SiT) with x-prediction v-loss flow matching to generate the target. This leverages spatial inductive biases for temporal structure.

  6. Downstream-utility evaluation protocol. The system is evaluated by training a classifier on real signals and scoring generated signals against true labels (train-on-real, test-on-synthetic), rather than point-wise reconstruction error. This measures label-relevant information transfer and avoids conflating fidelity with condition re-expression.

What the improved AI system can do:

  • Generate a missing physiological signal (e.g., continuous glucose monitor, mean arterial blood pressure) from multiple co-recorded, irregularly sampled, partially missing wearable or vital-sign modalities plus static clinical variables.

  • Maintain high downstream task utility (AUROC/AUPRC) across diverse clinical tasks—four-class study-group classification, sepsis, heart failure, and in-hospital mortality—even when baselines hover near chance.

  • Handle heterogeneous modalities with distinct dynamics and missingness patterns without treating them as homogeneous context.

  • Remain robust to sparse observations (e.g., 30% mask ratio) and to removal of either time-series or static conditioning groups, though both contribute.

  • Achieve performance at or above a real-signal reference on 13 of 16 evaluation settings, indicating label-relevant structure survives generation, sometimes in a cleaner form than the measured waveform.

  • Provide a stable, scalable framework that can be extended to jointly model cross-modal dependencies and separate fidelity from re-expression in future paired-data settings.

Abstract

Continuous physiological time series underpin modern clinical monitoring, yet many of the most informative signals are invasive, expensive, or simply unavailable for a given patient. Conditional generation offers a remedy: an absent signal can be synthesized from co-recorded signals and routine clinical variables. Existing generators, however, are built around a single conditioning modality and degrade when forced to handle the heterogeneous, irregularly missing mix of time-variant signals and static covariates seen in practice. We propose ReCoGen (Represent Conditions, then Generate), a two-stage framework that decouples multimodal condition representation from target generation. Stage I trains one masked autoencoder per modality, distilling each time-variant condition into a compact and missingness-tolerant token sequence. Stage II trains a flow-matching generator that fuses these tokens with static conditions to synthesize the target signal. Across three physiological benchmarks, including continuous glucose monitoring on AI-READI and arterial blood pressure generation on MIMIC-III and MIMIC-IV, ReCoGen attains the best downstream utility on all sixteen (dataset, task, metric) settings, surpassing six representative conditional generators; on thirteen of them its utility also reaches or exceeds the utility measured on the real signal, a reference we read as an approximate anchor rather than a ceiling. Ablations trace the gains to the conditioning path: learnable cross-attention over the frozen per-modality encoders, and a dual token-plus-AdaLN route for the static conditions. ReCoGen thus turns routinely collected signals into informative surrogates for invasive or unavailable ones, a step toward less invasive, lower-cost continuous clinical monitoring.

Sources

Related papers