Smooth Flow Matching for Synthesizing Functional Data
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Smooth Flow Matching for Synthesizing Functional Data".
Jane: The paper was written by Jianbin Tan and Anru R. Zhang from Duke University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're digging into a paper called "Smooth Flow Matching for Synthesizing Functional Data." Jane, I have to say, the title alone makes me want to understand what "functional data" actually means here.
Jane: Great question, Tom. Functional data is basically when your data points aren't single numbers but whole curves or trajectories. Think of a patient's blood pressure measured over time, or a stock price moving through a week. Each person gives you a function, not just a point.
Tom: So instead of saying "this patient's average blood pressure is one hundred twenty" you're saying "this patient's blood pressure follows this entire shape over time." That's a much richer way to look at things.
Jane: Exactly. And the challenge is, how do you generate new, realistic versions of these curves? That's what this paper tackles. The authors are Jianbin Tan and Anru Zhang from Duke University, and they've built a method that can create synthetic functional data that looks like the real thing.
Tom: And why would we want synthetic data at all? Why not just use the real data?
Jane: Because real medical data is sensitive. You can't just share patient records with researchers across the world. But if you can generate synthetic data that captures the same patterns without revealing any actual patient's information, you can share that instead.
Tom: So it's like creating a fake patient population that behaves like the real one, but isn't actually any real person. That's powerful for research and collaboration.
Jane: Right. And the "smooth" part of the title is important too. The curves they generate aren't jagged or noisy in a weird way—they're smooth, like real biological measurements would be.
Tom: I love that they're thinking about the actual nature of the data, not just throwing a generic model at it. So who's going to care about this? Hospitals? Researchers?
Jane: Anyone working with longitudinal data—medical records, climate measurements, even financial time series. The method is designed for situations where you have irregular, sparse observations, which is exactly what real-world data looks like.
Tom: Sparse and irregular, meaning patients get checked at different times, not on a perfect schedule. That's every hospital in the world.
Jane: Precisely. And that's what makes this paper exciting—it's solving a problem that's everywhere in practice.
Tom: Alright, I'm hooked. Let's get into what they actually did with this method.
Summary: Tom: So we've established that "Smooth Flow Matching" is about generating realistic synthetic curves from messy, real-world data. Jane, what's the core idea that makes this work?
Jane: The key insight is that they separate the problem into two parts. First, they figure out the distribution of values at each individual time point—so, what's the range of blood pressures at hour one, hour two, and so on. Then, separately, they figure out how those values are connected across time.
Tom: So it's like learning the personality of each moment, and then learning the story that connects those moments.
Jane: That's a nice way to put it, Tom. The technical term is a copula framework—you model the marginals and the dependence structure separately. This lets them handle data that isn't Gaussian, which is a big deal.
Tom: Why is non-Gaussian such a big deal? I feel like every paper I read mentions it.
Jane: Because a lot of older methods assume everything follows a bell curve. But real data—like heart rates or blood pressure—can be skewed, heavy-tailed, or have other weird shapes. If you force a bell curve on it, you lose important information.
Tom: And this method doesn't do that?
Jane: Correct. It learns the actual shape of the data at each time point without assuming it's Gaussian. That's a major advantage over the older, classical approaches.
Tom: Now, I saw they compared against some deep learning methods too. How does this stack up?
Jane: The deep learning methods—they call them DSM and FM—are powerful but they require densely observed data. They need to see the whole curve, or close to it, to learn. But real data is often sparse, with patients missing check-ins.
Tom: So those methods struggle when data is patchy.
Jane: Exactly. And that's where this paper shines. It's designed for sparse, irregular data from the ground up. In their simulations, it consistently matched or beat the other methods, and it was way faster to train.
Tom: Faster and better. That's a rare combination.
Jane: It is. And it's because they don't need a huge neural network to learn an operator between function spaces. They use a simpler, three-dimensional vector field, which is much more efficient.
Tom: So it's not just about accuracy—it's about practicality. You can actually run this on a regular computer.
Jane: Right. The paper shows computation times that are dramatically lower than the deep learning alternatives. That makes it accessible to more researchers.
Tom: I'm really impressed by the practical angle here. Let's talk about what this means for real-world applications.
Improvements: Tom: So we know this method works well in simulations. But what does it actually improve in the real world? Jane, you mentioned they tested it on medical data.
Jane: They did. They used the MIMIC-IV database, which is a huge repository of ICU patient records. They looked at blood pressure, respiratory rate, and heart rate over time.
Tom: And what did they find?
Jane: The synthetic data they generated with their method, "Smooth Flow Matching," looked remarkably like the real patient trajectories. The mean patterns and the main ways the curves vary—what they call eigenfunctions—matched up well.
Tom: So the fake data actually captures the real dynamics of patient health.
Jane: Exactly. But the other methods they compared against—the deep learning ones—produced data that looked quite different from the real thing. The curves were jagged, unrealistic, and didn't match the key patterns.
Tom: That's a huge difference. Why did the deep learning methods fail so badly here?
Jane: Because the real data is irregularly sampled. Patients get checked at random times, and there are gaps. The deep learning methods need dense, regular data, so the researchers had to interpolate—fill in the gaps—before training. That filling-in process introduced errors.
Tom: So garbage in, garbage out. The interpolation step corrupted the learning.
Jane: Right. And their method, "Smooth Flow Matching," doesn't need that step. It handles the irregularity directly, which is a major improvement.
Tom: They also mentioned privacy. How does that work?
Jane: They ran a privacy test to make sure the synthetic data wasn't just copying real patients. The results showed that the synthetic data was distinct from the training data, meaning it wasn't memorizing individual records.
Tom: So you get the utility of the data without the privacy risk. That's the holy grail for medical research.
Jane: It is. And they went a step further—they used the synthetic data to train a prediction model. The model trained on synthetic data performed almost as well as one trained on real data.
Tom: That's remarkable. You can train on fake data and still get real predictive power.
Jane: Exactly. And in some cases, it even did better, because the synthetic data was densely and regularly sampled, which reduced uncertainty.
Tom: So the improvement here isn't just about making pretty curves—it's about making useful data that can actually drive clinical decisions.
Jane: That's the real impact. It opens the door for sharing high-quality synthetic medical data across institutions without compromising patient privacy.
Tom: Let's bring in the rest of the team to get their take on this.
Conclusion: Tom: Alright, let's wrap this up. We've been discussing "Smooth Flow Matching for Synthesizing Functional Data," and I think we've only scratched the surface.
Jane: We really have. Let me bring in Lu and Meng to get their perspectives on the bigger picture.
Lu: Thanks, Tom. From a research standpoint, what excites me is that this method doesn't force data into a Gaussian box. It respects the true shape of the data, which is crucial for heavy-tailed distributions. That's a philosophical shift in how we approach functional data generation.
Meng: And from an engineering side, the computational efficiency is a game-changer. Training a deep neural operator can take hours and requires specialized hardware. This method runs in minutes on a standard machine. That means smaller labs and hospitals can actually use it.
Jane: That's a great point, Meng. It democratizes access to advanced generative modeling.
Lu: And the theoretical guarantees they provide—the consistency results—mean we can trust the method as the sample size grows. That's important for scientific credibility.
Meng: I also appreciate that they thought about the full pipeline. It's not just about generating data; it's about evaluating privacy and utility. They showed the synthetic data preserves predictive power, which is the ultimate test.
Tom: So we have a method that's accurate, fast, privacy-preserving, and theoretically sound. What's not to love?
Jane: The only limitation I see is that they focused on univariate functional data—one measurement at a time. Real EHR data often has multiple vital signs measured together, and capturing those interdependencies is the next frontier.
Lu: Absolutely. Extending this to multivariate functional data would be a natural and impactful next step. Imagine generating synthetic data that preserves the relationship between blood pressure and heart rate simultaneously.
Tom: That would be incredible for clinical research. You could simulate entire patient trajectories with all their complexity.
Meng: And it would make the synthetic data even more useful for training predictive models that rely on multiple inputs.
Jane: Well, that's a perfect segue to our next discussion. But before we go, let's give one final summary of "Smooth Flow Matching for Synthesizing Functional Data."
Tom: Please, Jane.
Jane: This paper introduces a method that generates realistic synthetic functional data—like patient trajectories—from sparse, irregular observations. It's fast, doesn't assume Gaussian distributions, and produces smooth, realistic curves. It outperforms deep learning methods on real medical data and offers strong privacy guarantees.
Tom: And it's a step toward making high-quality medical data shareable and usable for research without compromising patient privacy. That's a win for everyone.
Jane: Absolutely. Thanks for joining us, and we'll see you next time on the show.
Tom: Take care, everyone.
Jianbin Tan, Anru R. Zhang
Duke University
stat.ML, cs.LG
Submitted: 2026-08-10
Comments: Accepted for publication in The Annals of Statistics
Code: https://github.com/Jianbin-Tan/Smooth-Flow-Matching
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 72/100
The gist: The paper introduces a novel framework named Smooth Flow Matching (SFM), tailored for generative modeling of functional data that enables statistical analysis without exposing sensitive real data.
Key concepts
- Functional Data
- Data points that are not single numbers but entire curves or trajectories. Examples include tracking a patient's blood pressure over time or monitoring stock prices throughout a week, providing a rich shape rather than just an average point.
- Synthetic Functional Data
- Artificially generated curves that mimic the patterns and characteristics of real-world data. This allows researchers to share high-quality medical information for study without compromising the privacy of actual patient records.
- Copula Framework
- A technical method used to model data by separating two parts: the distribution of values at each individual time point (marginals) and how those values are connected across time (dependence structure). This handles non-Gaussian data effectively.
- Sparse, Irregular Observations
- Real-world data where measurements are taken at random times rather than on a perfect schedule. This is common in medical settings and presents a challenge for models that require dense, continuous data.
Terminology
Summary
The paper introduces a novel framework named Smooth Flow Matching (SFM), tailored for generative modeling of functional data that enables statistical analysis without exposing sensitive real data. Functional data, i.e., random functions observed over a continuous domain, are increasingly available in areas such as biomedical research, health informatics, and epidemiology. However, effective statistical analysis for functional data is often hindered by challenges such as privacy constraints, sparse and irregular sampling, infinite-dimensionality, and non-Gaussian structures.
Core Framework: Under a copula framework, SFM constructs a semiparametric smooth flow to generate infinite-dimensional functional data, free of Gaussianity and low-rank assumptions. It is computationally efficient, handles irregular observations, and guarantees the smoothness of the generated functions, offering a practical and flexible solution in scenarios where existing deep generative methods are not applicable.
Key Methodological Contributions:
-
Copula Functional Data: The paper defines a copula process with a base process, where the copula of X is determined by a latent base process Z. Proposition 1 establishes that a copula process can be generated by transporting a latent process through a family of continuous and strictly increasing functions.
-
Pointwise Flow for Functional Data: The authors propose to characterize both maps Ft−1 ◦ Ft,base and Ft,base−1 ◦ Ft within the framework of continuous normalizing flows. They construct a flow map φu,t such that φu,t(Z0(t)) = Zu(t), where Z0 is a base process and Z1 is the target function such that Z1(t) = X(t) in distribution. Theorem 1 establishes the existence and uniqueness of the solution, showing that φ1,t(·) = Ft−1 ◦ Ft,base(·), and that the inverse satisfies a pullback equation.
-
Smoothness Guarantees: Theorem 2 ensures that the solution to the flow equation lies in the Sobolev space Wq2(S) provided that the vector field satisfies certain smoothness conditions, guaranteeing that generated functions are smooth.
-
Pointwise Flow Matching: The paper generalizes flow matching to pointwise flow settings, solving an optimization problem to estimate the vector field. Theorem 3 shows that the minimizers obtained from the flow matching objective based on the functional sequence and from the conditional flow matching objective are the same.
-
Smooth Flow Matching for Discretely Observed Data: For irregularly and sparsely observed functional data, the paper proposes estimating the vector field via spline regression with a smoothness penalty, using the loss function Ln(U) based on observed data. The estimator V̂ is obtained by minimizing the penalized loss within a tensor product B-spline space.
-
Pullback Estimation and Forward Generation: The paper develops procedures to estimate the latent correlation function ρ for Gaussian copula processes, using the pullback equation to compute ψ̂1,t, then applying surface smoothing to estimate ρ̂. The complete SFM generation procedure is summarized in Algorithm 1.
Statistical Consistency: Theorem 4 establishes the consistency of the estimated vector field, showing that R(V̂, V) = op(1) as n → ∞, without requiring the number of observations per subject to grow. Theorem 5 establishes the consistency of the latent correlation function and generated functional data, with the Wasserstein distance W2(X, X̃) = op(1) as n → ∞.
Simulation Studies: Extensive simulation studies compare SFM with existing methods including DSM (Denoising Score Matching), FM (Flow Matching), GP (Gaussian Process Sampling), and KL (KL Expansion with FPCA). The results show that SFM consistently outperforms existing methods in both Gaussian and non-Gaussian (Gamma) settings while substantially reducing computational cost compared to deep generative methods. The paper states: SFM consistently produces the most accurate or at least competitive functional samples compared to all other methods.
Real Data Analysis: The paper applies SFM to clinical trajectory data from the MIMIC-IV patient electronic health records (EHR) longitudinal database, focusing on three clinical features: arterial blood pressure, respiratory rate, and heart rate. The analysis demonstrates that SFM effectively recovers the dominant longitudinal patterns and key distributional features of the original EHR data. The paper further assesses the synthetic performance of SFM through a privacy risk calculation and a prediction task, showing that SFM-generated data exhibit superior predictive power under privacy constraints. The results show that SFM achieves the lowest Errorg among all generative methods, indicating that it produces synthetic data with superior predictive power.
Advantages of SFM: The paper highlights several advantages over existing deep generative models and Gaussian-based methods: (1) SFM leverages flow-based techniques to model marginal distributions rather than full joint distributions, enabling direct flow estimation from irregular and sparse functional data; (2) SFM naturally incorporates smoothness into flow estimation, allowing it to borrow strength across curves and generate smooth functional data from sparse samples; (3) SFM is computationally efficient, with training times substantially shorter than DSM and FM (e.g., 30.28 seconds vs. 410.90 seconds for arterial blood pressure data).
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:
1. New Generative Model for Functional Data (SFM)
-
Implement the Smooth Flow Matching (SFM) framework as described in Algorithm 1 (Gaussian copula) and Algorithm S1 (Student-t copula)
-
Replace high-dimensional operator-based neural networks with a three-dimensional vector field (u, t, x) learned via penalized tensor-product B-splines (Equation 21)
-
Use the rectified flow construction (Equation 18) for conditional probability paths
-
Estimate the vector field via the empirical loss in Equation 22 with Monte Carlo sampling (H=1000 base samples, F=30 time grid points)
2. Handling Sparse and Irregular Observations
-
Process irregularly sampled functional data directly without requiring dense or complete observations
-
Apply the denoising procedure from Remark 4 (smoothing splines) before SFM training when noise is present
-
Use the pullback equation (Equation 23) to estimate the inverse flow and latent correlation function (Equation 25)
3. Smoothness Guarantee
-
Enforce smoothness of generated functions by restricting the vector field to cubic spline spaces with roughness penalties (J(U) in Equation 21)
-
Use Runge-Kutta integration (Proposition S1) with step size 0.01 to ensure generated functions remain in W22(S)
4. Copula-Based Semiparametric Generation
-
For Gaussian copula: estimate latent correlation via Equation 25, apply NPD approximation, generate via Algorithm 1
-
For Student-t copula: estimate degrees of freedom via composite pairwise likelihood (Part B.3), use Algorithm S1
-
For general copulas: use Algorithm S2 with maximum likelihood estimation of copula parameters
1. Generate Realistic Synthetic Functional Data
-
Produce smooth, non-Gaussian, heavy-tailed functional samples from sparse irregular observations
-
Match the true distribution in Wasserstein distance (as demonstrated in Figure 3, SFM outperforms DSM, FM, GP, and KL methods)
-
Recover key distributional features beyond mean and covariance (e.g., median function, as shown in Table S3)
2. Enable Privacy-Preserving Data Sharing
-
Generate surrogate data that preserves temporal dynamics without exposing sensitive real data
-
Achieve privacy loss near zero (Figure 7A), indicating no memorization of training data
-
Support downstream tasks like prediction with accuracy comparable to using real data (Figure 7B)
3. Handle Real-World EHR Data
-
Process irregularly sampled clinical trajectories (arterial blood pressure, respiratory rate, heart rate) from MIMIC-IV
-
Generate synthetic trajectories that recover dominant longitudinal patterns (Figure S17)
-
Achieve 5-10x faster training than deep generative methods (Table 2: SFM 30-55 sec vs DSM 410-598 sec)
4. Provide Statistical Guarantees
-
Consistent vector field estimation (Theorem 4) without requiring dense observations per subject
-
Consistent latent correlation estimation and Wasserstein-consistent generation (Theorem 5)
-
Guaranteed smoothness of generated functions (Theorem 2)
5. Extend to Various Copula Structures
-
Gaussian copula (Algorithm 1) for standard dependencies
-
Student-t copula (Algorithm S1) for tail dependence
-
General copula processes (Algorithm S2) for flexible dependence modeling
6. Support Predictive Modeling
-
Train regression models on synthetic data for one-step-ahead prediction (Equation 28)
-
Achieve prediction errors comparable to models trained on observed data, even with smaller sample sizes (Figure 7B)
Abstract
Functional data, i.e., random functions observed over a continuous domain, are increasingly available in areas such as biomedical research, health informatics, and epidemiology. However, effective statistical analysis for functional data is often hindered by challenges such as privacy constraints, sparse and irregular sampling, infinite-dimensionality, and non-Gaussian structures. To address these challenges, we introduce a novel framework named Smooth Flow Matching (SFM), tailored for generative modeling of functional data that enables statistical analysis without exposing sensitive real data. Under a copula framework, SFM constructs a semiparametric smooth flow to generate infinite-dimensional functional data, free of Gaussianity and low-rank assumptions. It is computationally efficient, handles irregular observations, and guarantees the smoothness of the generated functions, offering a practical and flexible solution in scenarios where existing deep generative methods are not applicable. Through extensive simulation studies, we demonstrate the advantages of SFM in terms of both synthetic data quality and computational efficiency. We then apply SFM to generate clinical trajectory data from the MIMIC-IV patient electronic health records (EHR) longitudinal database. Our analysis showcases the ability of SFM to produce high-quality surrogate data for downstream tasks, highlighting its potential to boost the utility of EHR data for clinical applications.
Sources
- Convergence of Continuous Normalizing Flows for Learning Probability Distributions
- Score Matching With Missing Data
- Sparse Equation Matching: A Derivative-Free Learning for General-Order Dynamical Systems
- Fourier Neural Operator for Parametric Partial Differential Equations
- Score-based Diffusion Models in Function Space
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Statistical Properties of Rectified Flow
- Boosting Data Analytics With Synthetic Volume Expansion
- Denoising Diffused Embeddings: a Generative Approach for Hypergraphs
- Functional Principal Component Analysis for Distribution-Valued Processes
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey