Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations

arXiv:2608.13234 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Gaute Johannessen, Geert Roelof van der Ploeg, Evrim Acar

University of Oslo · Simula Metropolitan Center for Digital Engineering

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/GauteJ1/CoupledModelsProject

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper introduces a knowledge-guided approach for pattern discovery that brings together data and computational models by jointly analyzing real data and simulated data (generated using a

Terminology

Summary

This paper introduces a knowledge-guided approach for pattern discovery that brings together data and computational models by jointly analyzing real data and simulated data (generated using a computational model) using coupled tensor factorizations with linear coupling. The approach is demonstrated on real metabolomics measurements, showing that guiding the analysis of noisy data with simulated data improves pattern discovery performance while also revealing potential discrepancies between data and computational models.

The paper states: "In this paper, we introduce a knowledge-guided approach that brings together data and computational models by jointly analyzing real data and simulated data (generated using a computational model) using coupled tensor factorizations with linear coupling. Our experiments on real metabolomics measurements demonstrate that guiding the analysis of such noisy data with simulated data improves the pattern discovery performance while also revealing potential discrepancies between data and computational models."

The method involves arranging real observational data as a higher-order tensor (e.g., metabolites by time by subjects) and simulated data as another tensor (e.g., metabolites by time by virtual subjects), then jointly analyzing them coupled in the features (metabolites) mode. The paper uses coupled tensor factorizations with linear coupling, specifically considering both CP and PARAFAC2 models for the real data. The coupling relation links factor matrices in the coupled mode via transformation matrices, allowing for partially overlapping feature sets.

The experiments use two metabolomics datasets: a real dataset from a meal challenge study (COPSAC2000 cohort) arranged as a 161 metabolites × 8 time points × 133 subjects tensor, and a simulated dataset generated using a human whole-body metabolic model, containing six blood metabolites/hormones (insulin, glucose, pyruvate, lactate, alanine, β-hydroxybutyrate) arranged as a 6 metabolites × 8 time points × 50 virtual subjects tensor.

The paper compares three settings: (1) CP (T0-corrected) vs. Linearly Coupled CP-CP (T0-corrected), (2) CP vs. Linearly Coupled CP-CP (without T0-correction, with nonnegativity constraints), and (3) PARAFAC2 vs. Linearly Coupled PARAFAC2-CP. Performance is assessed based on correlations between factors extracted from the subjects mode and body composition and insulin resistance measures (including HOMA-IR, muscle to fat ratio, fat percentage, BMI, waist circumference, etc.).

The results show that "jointly analyzing noisy real data with simulated data improves the recovery of a biologically meaningful component. The clean simulated data has an underlying component modeling co-varying insulin and glucose levels. When coupled with real data, the simulated data guides the analysis and enables recovery of this component known to be strongly linked to a BMI-related phenotype. In uncoupled models of real data, this component is either not captured because other metabolites dominate the model or is noisier. Therefore, the patterns extracted by the coupled models have consistently stronger correlations with body composition and insulin resistance measures."

Specifically, in the CP (T0-corrected) setting, a 2-component CP model of real data reveals a BMI-associated pattern dominated by metabolites other than insulin and glucose, while a 3-component linearly coupled CP-CP model forces a pattern mainly capturing insulin and glucose into the model, yielding stronger BMI-associated correlations. In the CP (nonnegative) setting, a 6-component CP model recovers an insulin-dominated factor but with a small glucose coefficient, while the coupled model finetunes this component, increasing the glucose coefficient and suppressing other metabolites, resulting in a purer insulin/glucose component with stronger correlations. Similarly, in the PARAFAC2 setting, the 6-component PARAFAC2 model recovers an insulin-dominated factor with a small glucose coefficient, while the coupled PARAFAC2-CP model captures a cleaner insulin/glucose pattern with stronger correlations.

The paper also notes a potential discrepancy: "the simulated time profile of this component shows a sharper peak around 1 hour post-meal than observed in the real data, where the time pattern is much broader using CP and shows much more individual variation using PARAFAC2. This demonstrates a potential discrepancy between the real data and the computational model. Furthermore, coupled models (via the factors extracted from the subjects mode) show that there is much less individual variation in the simulated data compared to real data."

The conclusion states: "we have introduced a knowledge guided approach for interpretable pattern discovery from complex data. We demonstrate that coupled tensor factorizations with linear coupling constraints are effective tools for bringing together data and computational models via joint analysis of real and simulated data. Our experiments on metabolomics data demonstrate that the proposed approach consistently provides cleaner patterns compared to the analysis of only real data. Besides, our results show that the coupled framework can reveal potential mismatches between real data and computational models." Future work is suggested on whether conflicting information in simulated and real data can be reliably extracted via coupled tensor factorizations (e.g., through unshared factors), and applying the approach in neuroscience.

Improvements for AI systems

Improvements to AI systems:

  1. Knowledge-guided tensor factorization for noisy data analysis – Enhance AI systems to jointly factorize real and simulated data tensors with linear coupling constraints, enabling the system to recover biologically meaningful patterns that are otherwise obscured by noise or dominant unrelated features in real data.

  2. Discrepancy detection between data and models – Build AI systems that automatically compare the time profiles and variability of coupled factors (e.g., sharp simulated peaks vs. broader real-data patterns) to flag mismatches between computational models and empirical observations, providing actionable feedback for model refinement.

  3. Coupled CP and PARAFAC2 with partial feature overlap – Implement AI systems that support flexible coupling via transformation matrices, allowing for partially overlapping feature sets (e.g., 6 simulated metabolites vs. 161 real metabolites), thus enabling transfer of clean structural knowledge from simulations to noisy high-dimensional real data.

  4. Guided component purification – Develop AI systems that use simulated data to clean up extracted factors, such as increasing the glucose coefficient in an insulin-dominated component and suppressing irrelevant metabolites, resulting in purer, more interpretable patterns with stronger correlations to external phenotypes (e.g., BMI, HOMA-IR).

  5. Multi-setting robustness – Create AI systems that can switch between T0-corrected, nonnegative, and PARAFAC2 formulations, automatically selecting the coupling strategy that yields the highest correlation with known biological outcomes, thereby improving generalizability across different data types.

  6. Individual variation quantification – Enable AI systems to quantify and compare subject-mode factor distributions between real and simulated data, revealing over-simplification in simulations (e.g., less individual variation) and guiding the generation of more realistic synthetic data.

What the improved AI system can do:

  • Analyze high-dimensional, noisy biomedical data (e.g., metabolomics, genomics) by integrating prior knowledge from mechanistic computational models, leading to more reliable discovery of disease-relevant patterns.

  • Automatically detect and report where simulations diverge from reality, helping researchers iteratively improve their models.

  • Recover subtle but biologically critical components (e.g., insulin-glucose co-variation) that standard uncoupled analysis misses, even when those components are weak relative to other signals.

  • Provide interpretable, phenotype-correlated factors (e.g., with body composition and insulin resistance) directly usable for biomarker discovery or patient stratification.

  • Adapt to different tensor structures (CP, PARAFAC2) and constraints (nonnegativity, T0-correction) without manual tuning, making it applicable to diverse temporal and longitudinal datasets.

  • Generate insights into inter-individual variability, distinguishing real biological heterogeneity from simulation artifacts, thus improving the fidelity of digital twins or virtual patient cohorts.

Sources

Related papers