Synthetic Data Generation for Augmenting Small Samples
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Synthetic Data Generation for Augmenting Small Samples".
Jane: The paper was written by Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram, Thomas Walters et al. from CHEO Research Institute and University of Ottawa and Charité - Universitaetsmedizin Berlin and Hospital for Sick Children and Ottawa Hospital Research Institute and McMaster University and OpenSourceResearch.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the health AI community, and it's called "Synthetic Data Generation for Augmenting Small Samples." Jane, I know you've been excited about this one all week.
Jane: Oh, absolutely, Tom. And honestly, the title tells you exactly what the problem is. In health research, we have these tiny datasets all the time. Maybe a few hundred patients, maybe a few thousand. And if you try to train a machine learning model on that, it just doesn't generalize well to new patients.
Tom: Right, and that's the core issue. The paper points out that in the literature, the median is about twelve point five events per predictor variable, which is way below the two hundred events per variable that you'd want for stable models.
Jane: So the idea here is simple on the surface. If you don't have enough real data, why not generate synthetic data that looks like your real data? Add it to your training set, and suddenly your model has more examples to learn from.
Tom: But here's the catch, and I love this about the paper. It's not always beneficial. The authors ran simulations on thirteen large health datasets, and they found that augmentation helps in some situations and hurts in others.
Jane: And that's actually the most important finding, Tom. Because a lot of people in the field just assume more data is always better. But this paper shows you need to be strategic about it.
Tom: Exactly. They built a decision support model that tells you whether augmentation is worth trying for your specific dataset. It's like a screening tool before you commit the computational resources.
Jane: And the results on seven small real-world datasets were pretty dramatic. They saw relative improvements in AUC ranging from about four percent up to forty-three percent. That's not a small effect.
Tom: Forty-three percent, Jane. That's the difference between a model that's basically guessing and one that's actually clinically useful.
Jane: So the title really captures it. This is about making small samples work better, and the paper gives you a roadmap for when and how to do that.
Tom: And I think the key thing listeners should take away from the title is that this isn't just a theoretical exercise. These are real health datasets, real clinical problems, and real improvements in prediction.
Jane: Right, and we're going to dig into the methodology next, because the way they set up these simulations is really clever.
Tom: Stay with us, folks. We're just getting started.
Summary: Tom: So we're back with "Synthetic Data Generation for Augmenting Small Samples," and Jane, you mentioned the methodology was clever. Let's break down what they actually did.
Jane: Okay, so imagine you have a big dataset, like a million hospital records. They split it into a training set and a test set. Then they draw smaller samples from the training set, like samples of size twenty, fifty, a hundred, all the way up to fifty thousand.
Tom: And for each of those small samples, they use four different generative models to create synthetic data. We're talking sequential decision trees, Bayesian networks, a conditional GAN, and a variational autoencoder.
Jane: Right, and they add varying amounts of synthetic data to each base sample. So for a base sample of a hundred patients, they might add ten synthetic patients, or a hundred, or a thousand. Then they train a gradient boosted decision tree on the augmented data and test it on the held-out test set.
Tom: And the key metric is AUC, which measures how well the model distinguishes between, say, patients who die and patients who survive.
Jane: Exactly. And what they found is that augmentation helps the most when the base dataset is small, when the baseline AUC is low, and when the dataset has high-cardinality categorical variables.
Tom: High-cardinality, meaning variables with many different categories, like a diagnosis code with thousands of possible values.
Jane: Right. And it makes sense, because those high-cardinality variables create a huge space of possible combinations. A small sample only covers a tiny fraction of that space, so synthetic data can fill in the gaps.
Tom: But here's the thing that surprised me. No single generative model consistently outperformed the others. Sometimes the Bayesian network was best, sometimes the GAN, sometimes the autoencoder.
Jane: That's a really important finding, Tom. It means you can't just pick your favorite model and assume it'll work. You need to evaluate all of them on your specific dataset.
Tom: And that's why they built the decision support model. It uses four characteristics of your dataset, the sample size, the imbalance of the outcome, the degrees of freedom, and the baseline AUC, to predict whether augmentation will help.
Jane: And that model had an AUC of zero point seven seven in leave-one-out cross-validation. So it's reasonably good at telling you whether to bother with augmentation.
Tom: So the summary is, augmentation can help a lot, but only under the right conditions, and you need to evaluate multiple generative models to find the best one.
Jane: And that's the practical guidance that clinicians and researchers need. It's not just "generate synthetic data and hope for the best." It's a systematic process.
Tom: Now, the really interesting part is the diversity argument. The paper claims augmentation helps not just because you have more data, but because the synthetic data is more diverse.
Jane: And that's exactly what we're going to talk about next. How they proved that diversity, not just sample size, is what drives the improvement.
Tom: Don't go anywhere, listeners.
Improvements: Tom: Welcome back. We're still on "Synthetic Data Generation for Augmenting Small Samples," and Jane, you just teased the diversity argument. Let's get into it.
Jane: So the question they asked was, is augmentation helping because you have more data, or because the synthetic data is genuinely different from what you already have? They needed to separate those two effects.
Tom: And they did that with a clever comparison. They compared augmentation using generative models against augmentation using resampling, which is basically just copying existing records with replacement.
Jane: Right. Resampling increases your sample size, but it doesn't add any new information. It's the same records, just repeated. So if augmentation with generative models performs better than resampling, then the benefit must come from diversity, not just size.
Tom: And that's exactly what they found. The augmented datasets had higher diversity, and they also had higher AUC. The difference was statistically significant.
Jane: They even developed a metric to measure diversity. They trained an isolation forest on the base dataset, then used it to score how many records in the augmented dataset looked like outliers relative to the original data.
Tom: So a synthetic record that's very different from anything in the original dataset counts as diverse. And the more diverse records you have, the better your model generalizes.
Jane: Exactly. And this is a really important improvement over just thinking about sample size. It tells us that the quality of synthetic data matters, not just the quantity.
Tom: But here's where I want to push back a little. The paper also found that augmentation can hurt performance for large datasets. So diversity isn't always good.
Jane: Right, and that makes sense. If your base dataset already covers the space well, adding diverse synthetic records might just introduce noise. The model starts learning patterns that don't exist in the real population.
Tom: So it's a Goldilocks situation. Too little augmentation and you don't get the benefit. Too much and you degrade performance. You need to find the sweet spot.
Jane: And the paper provides a process for that. You evaluate multiple generative models, you try different amounts of synthetic data, and you pick the one that maximizes AUC on a validation set.
Tom: They also found that the optimal amount of augmentation varies wildly. For the Hot Flashes dataset, they added seven hundred twenty synthetic records to three hundred sixty real ones. For the Diabetic Retinopathy dataset, they added over eleven thousand.
Jane: So there's no one-size-fits-all answer. You have to do the search for each dataset.
Tom: And that's computationally expensive, which is a limitation they acknowledge. But the decision support model helps you avoid wasting time on datasets where augmentation won't help.
Jane: Right. And I think the diversity finding is the most impactful contribution here. It changes how we think about synthetic data. It's not just a way to pad your dataset. It's a way to expose your model to a wider range of patient presentations.
Tom: And that could be huge for rare diseases, where you might only have a hundred patients but the disease manifests in many different ways.
Jane: Absolutely. And that's the kind of impact we'll talk about in our conclusion.
Tom: Stick around for our final thoughts.
Conclusion: Tom: So we're wrapping up our discussion of "Synthetic Data Generation for Augmenting Small Samples." Jane, what's the big picture here?
Jane: The big picture is that small datasets don't have to be a dead end for machine learning in health. This paper shows that with the right approach, you can use synthetic data to meaningfully improve predictive performance.
Tom: And the key is knowing when to use it. The decision support model they built, with that logistic regression, tells you based on your dataset size, outcome balance, variable complexity, and baseline performance.
Jane: Right. And if augmentation is recommended, you evaluate multiple generative models and find the right amount of synthetic data to add. The improvements can be substantial, like that forty-three percent jump in the Colposcopy dataset.
Tom: But the most profound finding for me is the diversity argument. Augmentation isn't just about having more data. It's about having data that exposes your model to a wider range of patient presentations.
Jane: And that's what helps the model generalize to unseen patients. The resampling comparison was really elegant because it showed that just copying records doesn't give you the same benefit.
Tom: Now, there are limitations. The computational cost is real. And they didn't look at fairness, which is a big concern when you're generating synthetic data. If your original data has bias, the synthetic data will probably have that same bias.
Jane: That's a really important point, Tom. And the paper acknowledges it as future work. But for now, the message is clear. If you're working with a small health dataset, augmentation is worth considering, and this paper gives you the tools to decide.
Tom: And I think that's the lasting impact. It moves synthetic data from a niche technique to a practical tool that clinicians and researchers can actually use.
Jane: So goodbye to "Synthetic Data Generation for Augmenting Small Samples." It's been a great discussion, and I think this paper will influence a lot of future work in the field.
Tom: Absolutely. And next time, we'll be looking at a paper on federated learning for multi-site clinical trials. That should be another good one.
Jane: Thanks for listening, everyone. See you next time.
Dan Liu, Samer El Kababji, Nicholas Mitsakakis, Lisa Pilgram, Thomas Walters, Mark Clemons, Greg Pond, Alaa El-Hussuna, Khaled El Emam
CHEO Research Institute · University of Ottawa · Charité - Universitaetsmedizin Berlin · Hospital for Sick Children · Ottawa Hospital Research Institute · McMaster University · OpenSourceResearch
cs.LG, cs.AI, stat.ML
Submitted: 2025-01-30
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
Key concepts
- Synthetic Data Generation
- This involves creating artificial data points that mimic real patient records. It is used to increase the volume of training examples when real-world datasets are too small, helping machine learning models learn better.
- High-Cardinality Variables
- These are variables, such as diagnosis codes, that have a large number of possible values. Small samples often fail to cover all possible combinations within these variables; synthetic data helps fill those gaps.
- Augmentation
- The process of adding generated synthetic records to an existing dataset. The goal is to increase the sample size and, if successful, improve the model's ability to generalize and predict outcomes accurately.
Terminology
Summary
Summary
This paper evaluates the use of data augmentation with generative models to improve the predictive performance of machine learning models trained on small tabular health datasets. The study addresses the common problem of small datasets in health research, which lead to suboptimal generalization performance of machine learning models. The authors state: To address this data scarcity problem, there is a growing interest in using data augmentation to simulate additional observations from existing data
and that augmentation can also be seen as a form of regularization, where the simulated data increase the diversity of the original dataset by generating more and different examples from the same population.
The study has two main contributions. First, it evaluates the augmentation performance of four synthetic data generation (SDG) methods—sequential decision trees (SEQ), Bayesian networks (BN), conditional tabular GAN (CTGAN), and tabular VAE (TVAE)—on the predictive performance of binary gradient boosted decision tree (GBDT) models using LightGBM. This evaluation was performed on 13 large, heterogeneous health datasets. The authors also developed a decision support model to recommend whether augmentation would be beneficial for a given dataset. Second, the study tests the benefit of data diversity by comparing augmentation with resampling (bootstrap) on seven small real-world datasets.
The simulation methodology involved splitting each large population dataset into 70% training and 30% test sets. From the training data, 40 base datasets of varying sizes (n0 from 20 to 50,000) were drawn using outcome-stratified random sampling. Each base dataset was augmented with synthetic data of various sizes (n'), creating a total of 12,000 augmented datasets per original dataset. The performance of the LightGBM models trained on these augmented datasets was evaluated on the held-out test set using the Area Under the Curve (AUC).
Key findings from the simulation include:
-
Augmentation improves prognostic performance for datasets that
have fewer observations, with smaller baseline AUC, have higher cardinality categorical variables, and have more balanced outcome variables.
-
No single generative model consistently outperformed the others: "Different generative models perform best depending on the dataset itself and its baseline size. Therefore, it is not possible to a priori say that a particular generative model is consistently superior for the augmentation task."
-
The benefits of augmentation diminish or become detrimental for large datasets, as
for a large base dataset, the increase in size has less marginal prognostic benefit and the dataset may already contain sufficiently diverse information.
The decision support model, a logistic regression, was trained to predict whether augmentation would be beneficial. It achieved an AUC of 0.77 and an accuracy of 76.15% in leave-one-out cross-validation. The model's key predictors are the baseline AUC (with the largest negative impact), the base dataset size (n0), the imbalance factor, and the degrees of freedom of the predictors. The authors note: The LR model shows that the baseline AUC has the biggest impact on whether to recommend augmentation, with lower baseline AUC baseline datasets benefiting more from augmentation.
For the second part of the study, the decision model recommended augmentation for all seven small application datasets. The results showed that augmentation led to a relative improvement in AUC ranging from 4.31% (AUC from 0.71 to 0.75) to 43.23% (AUC from 0.51 to 0.73), with an average relative improvement of 15.55%. This improvement was statistically significant (p=0.0078). The best generative model varied by dataset (e.g., CTGAN for Hot Flashes, TVAE for DCCG, BN for Breast Cancer Coimbra, CTGAN for Breast Cancer and Colposcopy/Schiller, BN for Diabetic Retinopathy, and TVAE for Thoracic Surgery).
The comparison between augmentation and resampling showed that Augmentation AUC was higher than resampling only AUC (p=0.016).
Furthermore, the diversity of augmented datasets, measured using a novel metric based on extended isolation forest outlier scores, was significantly higher than the diversity of resampled datasets (p=0.046). This supports the argument that the benefits of augmentation are due to increased data diversity, not just larger sample sizes. The authors state: Diversifying the existing data through augmentation plays an important role in enhancing the model performance, and therefore, increasing the sample size without making the data more diverse is not as beneficial.
The paper concludes with practical recommendations: "For datasets where the baseline AUC is high, augmentation may not provide a significant advantage. However, where the baseline AUC is medium or small, and where datasets are in the 100 to 3000 observations range, augmentation can potentially improve the performance of a model's AUC, sometimes by a considerable amount." The authors recommend using the decision support model to determine if augmentation is necessary and, if so, evaluating multiple generative models and degrees of augmentation to find the optimal level. They also emphasize the importance of avoiding data leakage by training generative models only on training data.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Integrate a pre-processing layer that automatically determines whether data augmentation will benefit model performance before training begins.
What it does:
-
Calculates four dataset characteristics: base sample size (n0), imbalance factor, degrees of freedom (accounting for categorical cardinality), and baseline AUC
-
Applies the logistic regression decision model (coefficients from Table 3) to output a probability of augmentation benefit
-
If probability > 0.5, proceeds with augmentation; otherwise, skips augmentation to save computational resources
Specific implementation:
augmentation score = 6.75 + (-4.79e-5 × n0) + (-4.94e-2 × imbalance factor)
+ (5.12e-4 × degrees of freedom) + (-7.63 × baseline AUC)
Improvement: Replace single-generator augmentation with an ensemble approach that evaluates four generative models (Sequential Decision Trees, Bayesian Networks, CTGAN, TVAE) and selects the best performer for each dataset.
Improvement: Optimize augmentation specifically to maximize data diversity rather than just sample size.
Improvement: For datasets with 100-3,000 observations, automatically trigger augmentation to improve model performance.
Improvement: Implement augmentation within cross-validation frameworks without data leakage.
Improvement: Automatically determine the optimal amount of synthetic data to generate.
Improvement: For datasets with high-cardinality categorical variables, prioritize augmentation to explore more of the categorical space.
Improvement: For datasets with balanced outcomes, apply augmentation more aggressively.
Improvement: Adjust augmentation strategy based on baseline model performance.
Improvement: Reduce computational cost by using the decision model to skip unnecessary augmentation evaluations.
Abstract
Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution. Augmentation increases sample size and is seen as a form of regularization that increases the diversity of small datasets, leading them to perform better on unseen data. We found that augmentation improves prognostic performance for datasets that: have fewer observations, with smaller baseline AUC, have higher cardinality categorical variables, and have more balanced outcome variables. No specific generative model consistently outperformed the others. We developed a decision support model that can be used to inform analysts if augmentation would be useful. For seven small application datasets, augmenting the existing data results in an increase in AUC between 4.31% (AUC from 0.71 to 0.75) and 43.23% (AUC from 0.51 to 0.73), with an average 15.55% relative improvement, demonstrating the nontrivial impact of augmentation on small datasets (p=0.0078). Augmentation AUC was higher than resampling only AUC (p=0.016). The diversity of augmented datasets was higher than the diversity of resampled datasets (p=0.046).
Sources
- Time Series Data Augmentation for Deep Learning: A Survey
- Why Tabular Foundation Models Should Be a Research Priority
- Hyperparameter Optimization: Foundations, Algorithms, Best Practices and Open Challenges
- Automatic Exploration of Machine Learning Experiments on OpenML
- Synthcity: facilitating innovative use cases of synthetic data in different data modalities
- On the Usefulness of Synthetic Tabular Data Generation
- Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks