Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Synthetic data for ratemaking".
Jane: Synthetic data generation, particularly for actuarial ratemaking, is explored as a solution to data scarcity and privacy concerns by benchmarking imputation-based methods against deep generative models.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's look at the title again, "Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders." It immediately tells us we are comparing traditional statistical imputation techniques against more modern, deep learning generative approaches in the context of insurance pricing.
Jane: The authors listed are Yevhen Havrylenko, Meelis Käärik, and Artur Tuttar. They come from institutions that have a strong background in both actuarial science and mathematics and statistics at universities in Switzerland and Estonia.
Lu: Their expertise seems perfectly positioned for this comparison because they bridge the gap between the statistical rigor needed for actuarial science and the advanced machine learning techniques used to create synthetic data.
Meng: I wonder how their specific background influences their choice of which models to benchmark, especially when considering deployment constraints at a real insurance company.
Lalam: The fact that they're focusing on ratemaking specifically shows they are tackling a niche but critical area where data scarcity hits hardest for operational modeling. It’s focused research with very real-world consequences for risk management systems.
The paper's summary: Tom: So, what does the paper actually conclude about these different generation methods? Basically, they are testing how well synthetic data preserves the original distributions of variables and the complex relationships between them when used in pricing models.
Jane: They found that Multivariate Imputation by Chained Equations with Random Forests is a very competitive approach for creating this high-fidelity tabular data compared to more complex generative models like adversarial networks.
Lu: That's interesting because they are highlighting the ease of use of MICE with Random Forests, suggesting it’s a strong contender against methods that usually require more intricate setup.
Meng: So, from a deployment standpoint, if we have existing R packages for MICE and Random Forests, this suggests we don't need to overhaul our entire data science infrastructure just to generate synthetic training sets.
Lalam: This finding is significant because it points towards a more accessible path for generating realistic datasets without needing the enormous computational overhead often associated with training large GANs or VAEs from scratch.
The paper's improvements: Tom: The authors suggest that the main improvement lies in choosing methods that balance statistical fidelity—how well the synthetic data mimics real distributions—with practical usability for actuaries who need to use the results in their work.
Jane: They specifically point out that MICE-based models, particularly MICE PART SYN and MICE FULL SYN, are identified as the best-performing methods overall across most of the metrics they tested.
Lu: They emphasize that this competitive edge for MICE is largely due to its simplicity and how well Random Forests handle different variable types without requiring extensive pre-processing steps.
Meng: That practical aspect is huge; if actuaries can implement this using existing tools in R, the barrier to entry for using synthetic data drops significantly, which impacts how quickly new models can be tested.
Lalam: This usability focus is key; it means that the advancement isn't just about getting a theoretically good result on a chart, but about giving the end-user—the actuary—a tool they can actually use in their daily work.
Conclusion: Tom: To wrap up, the main point of this paper is that for generating synthetic ratemaking data, MICE-based imputation methods are competitive with deep generative models when you factor in ease of use and the resulting statistical accuracy.
Jane: They conclude that MICE PART SYN and MICE FULL SYN stand out as the best performing methods overall based on their evaluation metrics for preserving data distributions and improving model coefficient estimation.
Lu: The implication here is that we can achieve high fidelity synthetic data without always resorting to the most computationally demanding deep learning architectures, opening up more avenues for research.
Meng: For me, it means we can move away from highly specialized training pipelines and towards more robust, reproducible statistical methods when augmenting our datasets for testing new pricing algorithms.
Lalam: This work suggests that making high-quality synthetic data generation accessible to actuaries is a much more achievable goal than relying solely on the most complex AI architectures.
Tom: We've covered a lot today about this paper, "Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders." It really shows that statistical methods can be incredibly effective when paired with smart modeling choices.
Jane: It’s exciting to see how this research points towards a more practical and accessible way for the actuarial community to handle data scarcity.
Lu: We're looking forward to seeing how future work builds on this comparison between imputation and deep generative techniques in complex domains.
Meng: I'm ready for whatever the next paper is, as long as it has a clear path toward practical implementation in an operational setting.
Lalam: I think this paper lays a foundation for creating more accessible and reliable data augmentation tools that can genuinely support the culture of innovation in risk science.
Yevhen Havrylenko, Meelis K¨a¨arik, Artur Tuttar
Department of Actuarial Science, Faculty of Business and Economics, University of Lausanne · Institute of Mathematics and Statistics, Faculty of Science and Technology, University of Tartu
stat.ML, cs.LG, stat.AP
Submitted: 2025-09-02
Updated: 2026-09-28
Comments: 49 pages, 7 figures, 4 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: Synthetic data generation, particularly for actuarial ratemaking, is explored as a solution to data scarcity and privacy concerns by benchmarking imputation-based methods against deep generative
Key concepts
- Multivariate Imputation by Chained Equations (MICE)
- MICE creates synthetic data by treating missing values like a problem to be solved. It iteratively predicts missing values for each variable based on the observed values of other variables using models like Random Forests until the imputed data stabilizes. This method is based on statistical relationships within the original dataset.
- Conditional Tabular Generative Adversarial Networks (CTGAN)
- CTGAN uses a Generator and a Discriminator to create realistic synthetic tabular data. The generator learns to produce new records by sampling from the actual data frequencies, while the discriminator tries to distinguish real data from synthetic data. It incorporates advanced techniques like mode-specific normalization for better results.
- Autoencoders (AEs)
- Autoencoders are neural networks used here in two ways: deterministically to convert complex categorical variables into numerical vectors, or via Variational Autoencoders (VAEs) to learn a probabilistic representation of the data. These representations can then be used as input for other models like CTGAN to generate synthetic records.
- Actuarial Ratemaking Data
- This refers to the high-quality datasets needed by actuaries to accurately calculate insurance risks and set appropriate prices for policies. Access is often limited by privacy or cost, which motivates the use of synthetic data generation techniques.
Terminology
Summary
Synthetic data generation, particularly for actuarial ratemaking, is explored as a solution to data scarcity and privacy concerns by benchmarking imputation-based methods against deep generative models. The core finding highlights that Multivariate Imputation by Chained Equations (MICE) with Random Forests offers a competitive, user-friendly approach for creating high-fidelity tabular data compared to complex adversarial networks and autoencoders, especially when assessing the ease of use for pricing actuaries.
Literature Overview and Motivation
Actuarial ratemaking requires high-quality data to accurately quantify risks and maintain pricing models, yet access is often limited by cost and privacy concerns. To address this, synthetic data generation is explored as a solution to create datasets that preserve the original marginal distributions of variables and multivariate relationships among covariates. The paper investigates generative methods previously studied in actuarial literature, such as Conditional Tabular Generative Adversarial Networks (CTGANs) and Variational Autoencoders (VAEs), alongside Multivariate Imputation by Chained Equations (MICE).
Multivariate Imputation by Chained Equations (MICE)
MICE is presented as a method based on treating data creation as a missing data problem, drawing synthetic values from the posterior predictive distribution of the original data. The algorithm operates iteratively:
-
For every missing value in a dataset, impute it using a simple technique and treat these initial imputed values as “placeholders.”
-
For a variable Yj set its “placeholders” back to missing.
-
Regress the observed values of Yj on other variables using some imputation model (e.g., Random Forests).
-
Replace missing values of Yj with predictions from the imputation model.
-
Repeat Steps 2-4 for each j with missing data to complete one cycle, iterating until convergence or a maximal number of iterations is reached.
Conditional Tabular Generative Adversarial Networks (CTGAN)
A GAN consists of a generator (G) and a discriminator (D), where G learns to produce synthetic data that fools D, which evaluates authenticity. CTGAN extends this framework by incorporating innovations such as mode-specific normalization, conditional generation of data, and handling of mixed data.
Specifically, CTGAN generates individual records by randomly selecting a variable and then choosing a value based on the actual data frequency. The generator employs skip connections to support categorical variables.
Autoencoders (AEs)
Autoencoders are used in two primary ways:
-
Deterministic AEs are employed to transform high-cardinality categorical variables into low-dimensional numerical vectors before fitting them to GAN-based models, replacing one-hot encoding.
-
Variational Autoencoders (VAEs) learn a probabilistic representation of data points in a latent space, which can be used for synthetic data generation. The paper also discusses
Separate AEs
where a separate AE is trained for each categorical variable to reduce its dimensionality to be used as input for CTGAN.
Comparative Study and Results
The comparative study uses the freMTPL2freq dataset to evaluate 10 different approaches, including CTGAN, MICE-based methods (MICE PART SYN, MICE FULL SYN), and VAE JAMOTTON. Evaluation is conducted using two types of metrics:
-
Dataset metrics assess the similarity between the distribution of original and synthetic data by calculating ratios like MAE(x orig j) and MAP E(x orig j) for numeric variables, as well as pairwise correlations.
-
Model metrics assess the impact on GLM performance, using coefficient difference metrics M1 (measuring distance between true and estimated coefficients based on synthetic data) and M2 (a combined metric).
The main results indicate that the MICE-RF approach is a competitive method for generation of synthetic ratemaking data in comparison to deep generative models such as a VAE and a CTGAN,
primarily due to its ease of use, given the existing packages in R and the versatility of RFs in dealing with variables of different types without much pre-processing.
MICE PART SYN and MICE FULL SYN are identified as the best-performing methods overall across most performance metrics.
Usability Assessment
Subjective assessment ranks MICE-based methods as the most accessible, primarily due to their streamlined implementation in the R package mice, which requires minimal setup beyond standard installation.
In contrast, CTGAN-based methods require additional setup time and more extensive data preprocessing,
and custom implementations necessitate a higher degree of code adaptation and environment configuration.
The paper concludes that MICE is the most user-friendly approach for actuaries.
Data Augmentation Impact
The study assesses the impact of augmenting original data with synthetic data on GLM performance. The results show that "data augmentation does not generally improve the performance of GLMs trained on the real data augmented with the synthetic one.
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper for its potential to improve existing AI systems, particularly those in actuarial science and synthetic data generation. The core contribution is demonstrating that MICE-based imputation methods are competitive with deep generative models (like CTGANs) for generating high-fidelity tabular insurance data, often with lower implementation complexity.
Here are the specific improvements and what the improved AI system can do:
)Improvements to Existing AI Systems based on this Paper:
-
Improve Synthetic Data Generation Fidelity and Efficiency
-
Enhance Model Robustness in Coefficient Estimation
-
Enable
Out-of-the-Box
Deployment for Actuarial Models -
Optimize Handling of High-Cardinality Categorical Variables
)Specific Capabilities of the Improved AI System:
-
Improved Synthetic Data Fidelity and Efficiency: The system can generate synthetic tabular datasets (like claim counts) that preserve both the original marginal distributions of variables and the multivariate relationships among covariates with high accuracy, comparable to deep generative models.
-
Enhanced Model Robustness in Coefficient Estimation: By using MICE-based methods (specifically MICE PART SYN and MICE FULL SYN), the system produces synthetic data where Generalized Linear Models (GLMs) trained on this data yield coefficient estimates that are statistically closer to the true underlying parameters than those obtained from models trained on deep generative model outputs.
-
Enable
Out-of-the-Box
Deployment for Actuarial Models: The system can be deployed using established statistical packages (like R's 'mice' package) rather than requiring complex, custom training pipelines and fine-tuning of deep neural networks. This drastically reduces the time and expertise required by actuaries to generate high-quality synthetic datasets for testing new models or augmenting existing ones. -
Optimize Handling of High-Cardinality Categorical Variables: The system can utilize a hybrid approach (CTGAN WITH AE MICE) where Autoencoders are used to learn low-dimensional representations of categorical variables before they are fed into the CTGAN, and MICE is then used to impute numeric variables. This combination specifically addresses the known weaknesses of pure GANs in generating high-cardinality categorical data, leading to better fidelity for these complex features.
)Summary of Impact:
The resulting AI system will be a versatile, reliable data augmentation tool for actuarial science that offers a compelling trade-off: it achieves high statistical accuracy (fidelity) and good model performance (coefficient estimation) while maintaining high usability and low operational complexity.
Abstract
Actuarial ratemaking depends on high-quality data, yet access to such data is often limited by the cost of obtaining new data, privacy concerns, etc. In this paper, we explore synthetic-data generation as a potential solution to these issues. In addition to generative methods previously studied in the actuarial literature, we explore and benchmark another class of approaches based on Multivariate Imputation by Chained Equations (MICE). In a comparative study using an open-source dataset, MICE-based models are evaluated against other generative models like Variational Autoencoders and Conditional Tabular Generative Adversarial Networks. We assess how well synthetic data preserves the original marginal distributions of variables as well as the multivariate relationships among covariates. The consistency between Generalized Linear Models (GLMs) trained on synthetic data with GLMs trained on the original data is also investigated. Furthermore, we assess the ease of use of each generative approach and study the impact of generically augmenting original data with synthetic data on the estimation of GLMs for predicting claim counts. Our results highlight the potential of MICE-based methods in creating high-fidelity tabular data while offering lower implementation complexity compared to deep generative models.
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey