Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems
Haiteng Wang, Yunfei Zhu, Tao Wang, Yikang Li, Jiabao Dong, Xiaoge Zhang, Lei Ren
Beihang University · Hong Kong Polytechnic University
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: 27 pages; 5 main figures and 4 extended data figures
Code: https://github.com/camaramm/tennessee-eastman-profBraatz
Project page: https://wang-fujin.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: PhysDGM is a stepwise physics-embedded diffusion generative model for synthesizing time-series data that are consistent with the underlying physical laws of dynamical systems.
Terminology
Summary
PhysDGM is a stepwise physics-embedded diffusion generative model for synthesizing time-series data that are consistent with the underlying physical laws of dynamical systems. PhysDGM embeds physical laws directly into each reverse diffusion step of the generative process, ensuring trajectory-level physical consistency, rather than enforcing constraints only at the final output. A large-scale AI-synthetic dataset (4.4 million samples, 20× scale-up) constructed by PhysDGM demonstrates strong fidelity across 34 datasets spanning turbofan engines, aero-engines, batteries, and chemical processes. After incorporating the synthetic data, the downstream task performance substantially surpassed that using real data alone by 48% for remaining useful life prediction, 15% for health indicator estimation, 22% for state-of-health assessment, and 20% for fault diagnosis. Moreover, it requires 10–20× less training data than existing approaches, substantially reducing the high cost of data collection in dynamical systems. We further demonstrate PhysDGM’s potential in identifying early-stage faults in aero-engines by incorporating AI-synthesized data. In summary, PhysDGM provides a solid foundation for generating physically consistent industrial time-series, paving the way for expanding physics-guided AI into diverse data-scarce environments, including both industrial machinery and complex chemical reaction dynamics.
The framework integrates physics-embedded diffusion with distribution guidance, incorporating dynamic physical constraints and physical laws (e.g., degradation, trends) to ensure physically plausible outputs. Dynamic physical constraints are progressively applied across training stages (Diff-early, Diff-mid, Diff-late). Gradient-guided sampling refines generated data by embedding physical laws through gradient adjustments, leveraging a pretrained denoiser without retraining.
PhysDGM generates physically-consistent synthetic data. Validation of explicit physical constraints in the synthetically generated data was performed via an ablation study comparing PhysDGM against the unconstrained baseline (PhysDGMnon-phys). Four distinct constraint types were evaluated: degradation constraints, coupling constraints, fixed range constraints, and discrete constraints. Performance was assessed using Dynamic Time Warping (DTW) for measuring sequence similarity, Euclidean Distance (ED) to quantify dissimilarity at corresponding time steps, ContextFID for statistical fidelity, and Correlation Loss for examining temporal dependencies. PhysDGM demonstrated substantial improvements over PhysDGMnon-phys, reducing error metrics by an average of 25.8% (DTW), 24.7% (ED), 84.1% (ContextFID), and 73.3% (Correlation Loss) across four constraints. These results indicate that incorporating physical constraints significantly enhances the structural and statistical similarity of synthetic data to the original data.
Stepwise physical embedding significantly improves data quality. The efficacy of the stepwise physical embedding strategy was assessed by benchmarking it against a traditional PINN approach that incorporates physical constraints in the loss function of the neural network. Specifically, separate transformer models were trained using synthetic data generated from each method and their RMSE on the original test sets was compared. The results demonstrate that the data generated by PhysDGM shows significantly higher utility, reducing RMSE by 59.7%, 14.6%, 50.0%, and 12.6% across four datasets compared to the PINN baseline. This finding strongly suggests that the progressive constraint scheme achieves a superior balance between capturing complex data distributions and adhering to physical laws, thereby producing synthetic data with substantially higher physical significance.
PhysDGM improves equipment health prognostics. The practical engineering utility of PhysDGM was validated by applying it to three critical predictive maintenance tasks: turbofan engine RUL prediction, aero-engine HI prediction, and battery SOH estimation. In these tasks, original training data was augmented with PhysDGM-generated synthetic samples to train downstream prediction models at ratios ranging from 100% (1x) to 2000% (20x). Performance was systematically compared against models trained only with the original data and those augmented by samples generated by advanced generative models, such as DiffWave and DiT. As the volume of PhysDGM’s synthetic data increased, the RMSE for all prediction tasks showed a steady downward trend. Specifically, increasing the augmentation scale from 1x to 20x further reduced RMSE by 21.0%, 12.2%, and 11.1% for the RUL, HI, and SOH tasks, respectively. More importantly, compared to models trained solely on raw data, the 20x augmentation strategy reduced the RMSE for three prediction tasks by 47.6%, 15.4%, and 21.5%. Conversely, competing models such as DiT exhibited performance degradation at higher augmentation volumes. This observation suggests that scaling up low-fidelity synthetic data, which often lack physical plausibility, introduces noise or spurious patterns that mislead downstream models. In contrast, PhysDGM outperforms all baselines, indicating that it generates high-fidelity samples that effectively expand the training manifold and enhance predictive accuracy.
PhysDGM improves fault diagnosis in chemical processes. To systematically evaluate the effectiveness of PhysDGM for chemical process fault diagnosis, a benchmark chemical fault diagnosis dataset was used and models trained on real data alone were compared with those trained on 1:1 mixtures of real data and synthetic data from multiple generators. Performance was assessed using two widely adopted metrics—accuracy and area under the receiver operating characteristic curve (AUROC). It was found that augmenting real data with PhysDGM-generated samples consistently surpassed both the real-only baseline and mixtures with alternative generators. In terms of accuracy, PhysDGM-augmented model achieved gains of ↑13% over real-only, ↑117% over mixtures of real and DiT, and ↑131% over mixtures of real and DiffWave. For AUROC, the corresponding improvements were ↑4%, ↑18%, and ↑22%, respectively.
PhysDGM achieves full-data performance with 20× less training data. The capacity of PhysDGM to generate high-fidelity data under extreme data scarcity condition was assessed by training the model on a subset containing only 5% of the turbofan engine RUL prediction dataset. Next, downstream prediction performance was systematically compared across three training configurations: (1) a low-data baseline using only 5% original data, (2) a performance benchmark using 100% complete data, and (3) a data-augmented set combining the 5% original data with generated synthetic data. Remarkably, data augmentation restored prediction accuracy to levels comparable with the full dataset. Compared to the low-data baseline, the augmentation strategy yielded significant RMSE reductions, ranging from 29.97% to 47.56%. This finding indicates that PhysDGM can efficiently capture the intrinsic distribution and key features from extremely limited samples and generate high-fidelity data that effectively substitutes for large amounts of real data.
PhysDGM markedly enhances industrial prediction tasks across diverse models. Adding PhysDGM-generated synthetic data to the training set significantly improves the performance of diverse downstream models on industrial prediction tasks. RMSE and the coefficient of determination (R2) on the turbofan engine RUL prediction and battery SOH estimation datasets were used. The comparison includes models trained only on real data and models trained on a 1:1 mixture of real and synthetic data of different types. In terms of RMSE, replacing the original training data with synthetic data generated by PhysDGM reduces the RMSE of CNN by 7.93 (29.85 → 21.92), LSTM by 4.03 (25.29 → 21.26), Transformer by 9.01 (26.76 → 17.75), TLSTM by 0.75 (14.68 → 13.93), and MCTAN by 8.97 (27.33 → 18.36). In terms of 1 − R2, the reductions are 0.26 (0.55 → 0.29) for CNN, 0.11 (0.38 → 0.27) for LSTM, 0.31 (0.50 → 0.19) for Transformer, 0.02 (0.13 → 0.11) for TLSTM, and 0.27 (0.47 → 0.20) for MCTAN. By contrast, mixing real data with synthetic data generated by DiT or DiffWave yields improvements in these metrics that are no more than 5% of the gains achieved with PhysDGM.
Improvements for AI systems
Improvements to AI Systems:
- Physics-Embedded Diffusion for Time-Series Generation
-
Integrate stepwise physical constraints (degradation, coupling, fixed ranges, discreteness) directly into each reverse diffusion step, not just the final output.
-
Use gradient-guided sampling with a pretrained denoiser to enforce physical laws without retraining.
-
Resulting capability: Generate synthetic time-series data that are physically consistent at the trajectory level, reducing structural errors by 25% (DTW/ED) and statistical fidelity errors by 84% (ContextFID) compared to unconstrained generation.
- Progressive Constraint Scheduling (Diff-early, Diff-mid, Diff-late)
-
Apply dynamic physical constraints at different training stages to balance distribution learning and law adherence.
-
Resulting capability: Outperform PINN-based generation by reducing downstream RMSE by 12.6–59.7% across diverse datasets, enabling high-utility synthetic data for model training.
- Data Augmentation for Predictive Maintenance
-
Augment real training data with PhysDGM-generated samples at scales up to 20×.
-
Resulting capability: Improve remaining useful life prediction by 48%, health indicator estimation by 15%, state-of-health assessment by 22%, and fault diagnosis by 20% over real-data-only models, while avoiding performance degradation seen with DiT/DiffWave at high augmentation volumes.
- Extreme Data Scarcity Handling
-
Train the generator on only 5% of original data, then augment downstream training.
-
Resulting capability: Achieve full-data prediction accuracy with 20× less real data, reducing RMSE by 29.97–47.56% compared to low-data baselines—critical for expensive industrial data collection.
- Model-Agnostic Performance Enhancement
-
Use synthetic data to replace or augment training sets for diverse architectures (CNN, LSTM, Transformer, TLSTM, MCTAN).
-
Resulting capability: Reduce RMSE by up to 9.01 points and improve R2 by up to 0.31 across models, with gains 20× larger than those from DiT/DiffWave synthetic data.
- Early Fault Detection in Aero-Engines
-
Incorporate AI-synthesized data into diagnostic pipelines.
-
Resulting capability: Identify incipient faults earlier than real-data-only models, enabling proactive maintenance and reducing unplanned downtime.
- Cross-Domain Industrial Applicability
-
Apply the framework to turbofan engines, batteries, and chemical processes.
-
Resulting capability: Provide a generalizable solution for physics-guided AI in data-scarce environments, including complex chemical reaction dynamics, without task-specific retuning.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks