From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition

arXiv:2308.04553 · cs.CV, cs.LG · Submitted 2026-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition".

Jane: The paper was written by Maan Qraitem, Kate Saenko and Bryan A. Plummer from Boston University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone. Today we're digging into a paper that's got a great title, "From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition." Jane, I gotta say, that title alone got me hooked.

Jane: Oh, absolutely, Tom. And it's from Boston University — Maan Qraitem, Kate Saenko, and Bryan Plummer. The title basically tells you the whole story, right? They're using fake images to help models learn better on real ones.

Tom: So let's break that down for our listeners. What's a spurious correlation, anyway?

Jane: Think of it like this. You train a model to recognize dogs, but in your training photos, every big dog happens to be outdoors. The model might just learn "outdoors equals big dog" instead of actually learning what a big dog looks like. That's a spurious correlation — a shortcut the model takes.

Tom: And that shortcut can be a real problem in the real world. You show the model a big dog indoors, and it might completely miss it.

Jane: Exactly. And that's where the synthetic images come in. The idea is, you can generate fake images of big dogs indoors to balance out your training data. But here's the twist — the paper shows that just mixing fake and real images together creates a whole new problem.

Tom: Which is what? I'm guessing it's not as simple as just adding more data.

Jane: Right. Because fake images have their own quirks, their own artifacts. So the model might start using "this looks synthetic" as a shortcut, which is just another spurious correlation. The paper's whole point is to fix that.

Tom: So they're not just saying "use fake data." They're saying "here's the right way to use fake data." That's a big deal.

Jane: It really is. And it's one of those ideas that seems obvious once you hear it, but nobody had really laid it out this clearly before.

Summary: Tom: So we've got the title, we've got the problem. Let's get into what this paper actually does. Jane, can you walk us through the core idea?

Jane: Sure. The authors propose something called FFR — From Fake to Real. And it's a two-step training process. Step one, you pre-train a model on a balanced set of synthetic images. Step two, you fine-tune that model on the real data.

Tom: And that's it? Just two steps?

Jane: That's the beauty of it. It's simple. But the key is that you never mix the fake and real images in the same training batch. By keeping them separate, the model never gets the chance to learn that "synthetic" is a useful signal.

Tom: So the model learns the actual features from the fake data first, and then it adapts to the real data without ever seeing them side by side.

Jane: Exactly. And they have a theoretical proof showing that if you mix the data, you're guaranteed to introduce a bias toward whether an image is real or synthetic. It's mathematically unavoidable.

Tom: Wait, guaranteed? That's a strong claim.

Jane: It is, but they prove it. The bias in the original dataset is so baked in that any attempt to fix it with synthetic data just moves the bias somewhere else — specifically, to the real-versus-synthetic distinction.

Tom: So the solution is to not give the model that distinction at all. That's clever.

Jane: It's one of those things where you read it and think, "Why didn't I think of that?" But it takes a lot of careful analysis to get there.

Tom: And the results back it up. They tested this across three datasets and multiple bias levels, and FFR consistently beat the older methods.

Jane: By a lot, in some cases. Up to twenty percent improvement in worst-case accuracy. That's not a small bump.

Improvements: Tom: Okay, so we know FFR works. But what does it actually improve over? Jane, what were people doing before?

Jane: So there were two main approaches. One was called Additive Synthetic Balancing — you just add a bunch of balanced fake data to your real dataset. The other was Uniform Synthetic Balancing — you generate enough fake data to make every subgroup the same size.

Tom: And both of those have the problem we talked about — the model picks up on the real-versus-synthetic signal.

Jane: Right. And the paper shows that even if you combine those methods with other bias-mitigation techniques, like GroupDRO or resampling, they still underperform FFR.

Tom: Because those techniques are trying to fix a bias that FFR just avoids entirely.

Jane: Exactly. FFR doesn't need to fix the real-synthetic bias because it never lets it develop in the first place. And that means the other bias-mitigation methods can focus on the original bias — the one that actually matters.

Tom: So it's not just a new method. It's a better foundation for existing methods.

Jane: That's a great way to put it. They show that FFR improves the performance of GroupDRO, resampling, and even DFR — which is a pretty strong method on its own.

Tom: And they did all these ablations too, right? Testing different parts of the pipeline.

Jane: Yeah, they tested whether you even need the synthetic pretraining, and the answer is a clear yes. They also tested what happens if you pre-train on biased synthetic data instead of balanced data — and performance drops significantly.

Tom: So the balance in the synthetic data really matters.

Jane: It does. If your fake data has the same bias as your real data, you're not fixing anything. You're just reinforcing it.

First Page: Tom: Let's actually look at the first page of the paper, because there's a really nice figure there. Jane, can you describe it for our listeners?

Jane: Sure. There's a figure showing saliency maps — basically, heatmaps of where the model is looking when it makes a prediction. And they're comparing three approaches: the old synthetic augmentation methods, and then FFR.

Tom: And what do the heatmaps show?

Jane: For the old methods, the model is looking at all the wrong places. Background, furniture, even artifacts in the synthetic images — like a dog with three toes instead of four. It's using those weird details to make its prediction.

Tom: Yikes. So it's not even looking at the dog.

Jane: Not really. But with FFR, the model focuses right on the dog's face and body. It's learned the actual features that matter, not the shortcuts.

Tom: That's a really powerful visual demonstration. You can see the difference immediately.

Jane: And it connects directly to their theoretical point. The old methods are biased toward the synthetic-real distinction, so they use different features for fake images than for real ones. FFR uses the same features for both.

Tom: So the model is actually learning something generalizable.

Jane: Exactly. And they even show this with t-SNE plots — the real and synthetic images cluster together with FFR, but they form separate clusters with the old methods.

Tom: That's a nice confirmation that the model isn't treating them differently.

Jane: It is. It's one thing to say your method works. It's another to show that the model is actually learning the right thing.

Conclusion: Tom: Alright, we're wrapping up our discussion of "From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition." Jane, what's the big takeaway for our listeners?

Jane: I think the big takeaway is that synthetic data is a powerful tool, but you have to be careful about how you use it. Just throwing it into your training set can create new problems.

Tom: And FFR gives you a simple, clean way to avoid those problems.

Jane: Exactly. Pre-train on balanced fake data, fine-tune on real data, and you're done. It's easy to implement, it works with existing methods, and it gives you a significant boost in worst-case performance.

Tom: And that worst-case performance matters, right? That's the model's performance on the groups it struggles with the most.

Jane: Absolutely. That's where bias shows up. And improving that by up to twenty percent is a real win.

Tom: So, any downsides or open questions?

Jane: The authors mention that the generative model itself can have biases. If your fake data isn't actually balanced, FFR won't help as much. And they note that diffusion models sometimes struggle with complex prompts.

Tom: So the quality of your synthetic data still matters.

Jane: It does. But the framework itself is solid. And it's flexible — you can use it with any generative model and any bias-mitigation method.

Tom: Well, I think this one's going to be influential. It's a clean idea, well-proven, and easy to adopt.

Jane: I agree. And with synthetic data becoming more and more common, this kind of guidance is going to be essential.

Tom: Thanks for listening, everyone. We'll be back with the next paper soon.

Maan Qraitem, Kate Saenko, Bryan A. Plummer

Boston University

cs.CV, cs.LG

Submitted: 2026-08-08

Comments: Accepted at ECCV 2024

Code: https://github.com/mqraitem/From-Fake-to-Real

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

Terminology

Summary

Summary

This paper introduces a two-stage training pipeline called From Fake to Real (FFR) to mitigate spurious correlations in visual recognition models when using synthetic data augmentation. The authors identify a critical flaw in prior synthetic augmentation methods—Additive Synthetic Balancing (ASB) and Uniform Synthetic Balancing (USB)—which train on a mixed distribution of real and synthetic data. They argue that this mixed training introduces a new source of bias: the model can learn correlations between the bias variable B and the data source variable G (Real vs. Synthetic), i.e., the pair (B, G). For example, a model might learn that Synthetic Indoors predicts Big Dogs rather than learning the true class features.

The paper provides a theoretical proof (Theorem 1) showing that, under the standard spurious-correlation setting, every possible augmentation of a biased dataset with synthetic data will exhibit some bias toward (B, G). This is because the distributional differences between real and synthetic data (e.g., generative model artifacts) cannot be eliminated by simply balancing the data.

To address this, FFR separates the two data sources into two distinct training stages:

  • Stage 1: Pre-train a model on a balanced synthetic dataset where P D syn(YB) = P D syn(Y). This learns robust representations across all subgroups.

  • Stage 2: Fine-tune the model on real data using either Empirical Risk Minimization (ERM) or common loss-based bias mitigation methods (e.g., GroupDRO, Resampling, Deep Feature Reweighting).

By training on real and synthetic data separately, FFR prevents the model from being exposed to the statistical differences between the two data sources, thereby avoiding the bias toward (B, G).

The authors conduct extensive experiments over three datasets—CelebA-HQ, UTK-Face, and SpuCo Animals—across five bias ratios (90%, 95%, 97%, 99%, 99.9%). Key findings include:

  • FFR improves worst-group accuracy over prior synthetic augmentation methods (USB and ASB) by up to 20%.

  • FFR remains stable as bias severity increases, whereas prior methods degrade significantly.

  • Combining FFR with synthetic-data-free bias mitigation methods (GroupDRO, Resampling, DFR) yields further improvements, with FFR boosting their worst accuracy by 13%, 13%, and 4%, respectively.

  • Ablations confirm that both stages are necessary, and that balanced pretraining is crucial (biased pretraining drops performance significantly).

  • t-SNE visualizations show that FFR projects real and synthetic embeddings more tightly than USB or ASB, indicating it uses consistent core features rather than data-source-specific artifacts.

  • Qualitative saliency maps (RISE) show that FFR focuses on relevant object features (e.g., dog features) while prior methods attend to spurious background or synthetic artifacts.

The paper also includes supplementary analyses on lower-quality synthetic images, bias-agnostic data generation, hyperparameter details, and experiments with ViT backbones, all confirming FFR's robustness and superiority.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace single-stage mixed-data training with a two-stage approach:

  • Stage 1: Pre-train on balanced synthetic data (equal samples per bias subgroup) to learn robust, unbiased representations.

  • Stage 2: Fine-tune on real data using ERM or loss-based bias mitigation (GroupDRO, Resampling, DFR).

What the improved system can do:

  • Avoid learning spurious correlations between the target class and the pair (bias, data source) such as synthetic indoors → big dogs.

  • Achieve up to 20% higher worst-group accuracy compared to prior synthetic augmentation methods (USB, ASB) across CelebA-HQ, UTK-Face, and SpuCo Animals.

  • Maintain stable performance even at extreme bias ratios (99.9%) where prior methods degrade significantly.

Abstract

Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions B (, Indoors) are over-represented in certain classes Y (, Big Dogs). Synthetic data from off-the-shelf large-scale generative models offers a promising direction to mitigate this issue by augmenting underrepresented subgroups in the real dataset. However, by using a mixed distribution of real and synthetic data, we introduce another source of bias due to distributional differences between synthetic and real data (synthetic artifacts). As we will show, prior work's approach for using synthetic data to resolve the model's bias toward B do not correct the model's bias toward the pair (B, G), where G denotes whether the sample is real or synthetic. Thus, the model could simply learn signals based on the pair (B, G) (, Synthetic Indoors) to make predictions about Y (, Big Dogs). To address this issue, we propose a simple, easy-to-implement, two-step training pipeline that we call From Fake to Real (FFR). The first step of FFR pre-trains a model on balanced synthetic data to learn robust representations across subgroups. In the second step, FFR fine-tunes the model on real data using ERM or common loss-based bias mitigation methods. By training on real and synthetic data separately, FFR does not expose the model to the statistical differences between real and synthetic data and thus avoids the issue of bias toward the pair (B, G). Our experiments show that FFR improves worst group accuracy over the state-of-the-art by up to 20% over three datasets. Code available: https://github.com/mqraitem/From-Fake-to-Real

Sources

Related papers