A Consensus Privacy Metrics Framework for Synthetic Data

summary

Video file (mp4)

The gist

Synthetic data generation is one approach for sharing individual-level data.

In short

The episode reviews "A Consensus Privacy Metrics Framework for Synthetic Data," a global effort to standardize how synthetic data is measured for safety. The panel argues that traditional metrics like similarity and differential privacy's epsilon are inadequate. They propose focusing on membership and attribute disclosure, using quasi-identifiers and non-member baselines, providing a common language for safer synthetic data sharing.

Key concepts

Membership Disclosure
This measures whether an attacker can figure out if a specific person was part of the original training data used to generate the synthetic dataset. The framework suggests calculating this by comparing the F1 score against a naive guess, providing a quantifiable measure of leakage.
Attribute Disclosure
This measures whether an attacker can predict sensitive information, like a diagnosis, for an individual. It is determined by comparing how well the model predicts outcomes for people who were in the training set versus those who were not in the training set.
Differential Privacy (Epsilon)
A mathematical value used to quantify privacy protection. The panel notes that if epsilon is high, such as 4 or 18, the theoretical guarantee does not translate into actual practical privacy protection, requiring empirical testing of the synthetic data.
Quasi-Identifiers
These are attributes an attacker might know about a person, such as age or zip code. The framework recommends basing privacy metrics on these specific attributes rather than using every single attribute in the dataset, as this provides more relevant context for risk assessment.

Terminology used across episodes

This episode discusses

The paper

A Consensus Privacy Metrics Framework for Synthetic Data · Read on arXiv

Lisa Pilgram, Fida K. Dankar, Jorg Drechsler, Mark Elliot, Josep Domingo-Ferrer, Paul Francis, Murat Kantarcioglu, Linglong Kong, Bradley Malin, Krishnamurty Muralidhar, Puja Myles, Fabian Prasser, Jean Louis Raisaro, Chao Yan, Khaled El Emam

University of Ottawa · CHEO Research Institute · Charité – Universitätsmedizin Berlin · Institute for Employment Research · Ludwig-Maximilians-Universität · University of Maryland · University of Manchester · Universitat Rovira i Virgili · Max Planck Institute for Software Systems · Virginia Tech · University of Alberta · Vanderbilt University Medical Center · Vanderbilt University · University of Oklahoma · Medicines and Healthcare products Regulatory Agency · Berlin Institute of Health at Charité – Universitätsmedizin Berlin · University Hospital Lausanne

Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.

DOI: 10.1016/j.patter.2025.101320

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Consensus Privacy Metrics Framework for Synthetic Data".

Jane: The paper was written by Lisa Pilgram, Fida K. Dankar, Jorg Drechsler, Mark Elliot, Josep Domingo-Ferrer et al. from University of Ottawa and CHEO Research Institute and Charité – Universitätsmedizin Berlin and Institute for Employment Research and Ludwig-Maximilians-Universität and University of Maryland and University of Manchester and Universitat Rovira i Virgili and Max Planck Institute for Software Systems and Virginia Tech and University of Alberta and Vanderbilt University Medical Center and Vanderbilt University and University of Oklahoma and Medicines and Healthcare products Regulatory Agency and Berlin Institute of Health at Charité – Universitätsmedizin Berlin and University Hospital Lausanne.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, we’ve got a paper today that’s going to make a lot of people in data science sit up straight. It’s called “A Consensus Privacy Metrics Framework for Synthetic Data,” and honestly, the title alone tells you we’re finally getting serious about how we measure privacy.

Jane: Absolutely, Tom. And I love that this isn’t just one lab’s opinion. You’ve got researchers from universities, government agencies, and even regulators across Canada, Germany, the UK, Spain, the US, Switzerland — a real global panel. They ran a Delphi process, which is basically a structured way to get experts to agree on something through multiple rounds of voting.

Tom: Right, and that’s huge because synthetic data is being shared everywhere now — health records, census data, clinical trials. But until now, there wasn’t a standard way to say, “Yes, this synthetic dataset is safe to release.”

Jane: Exactly. And the paper’s main message is pretty bold: a lot of the metrics people currently use to claim privacy are either misleading or just wrong. They specifically call out similarity metrics — you know, how close a synthetic record is to a real one — and say those shouldn’t be used as standalone privacy guarantees.

Tom: That’s a big deal because those are the most common metrics in the literature. I mean, if you’re a company selling synthetic data and you’re saying, “Look, our data is private because no synthetic row exactly matches a real row,” this paper is basically saying that’s not enough.

Jane: Right, because similarity doesn’t capture what actually matters — whether someone can figure out if a person was in the training data, or infer something sensitive about them. That’s what they call membership disclosure and attribute disclosure.

Tom: And those are the two types of disclosure the panel says we should focus on. They’re not saying synthetic data is unsafe, just that we need better tools to measure it.

Jane: And that’s what the rest of the paper is about — giving us those tools, or at least the framework to build them. I’m excited to dig into the details.

Tom: Me too. Let’s get into the summary and see what the panel actually recommends.

Summary: Jane: So, Tom, we’re back with “A Consensus Privacy Metrics Framework for Synthetic Data,” and the summary really lays out the core findings. The panel basically says we need to stop relying on similarity metrics and the privacy budget epsilon from differential privacy as the main ways to report privacy.

Tom: Yeah, and that epsilon point is spicy. For people who don’t know, differential privacy gives you a number, epsilon, that’s supposed to tell you how much privacy you’re getting. Lower is better. But the panel says unless epsilon is really close to zero, that number doesn’t mean much in practice.

Lu: And that’s a really important nuance, Tom. A lot of real-world systems use epsilon values like four eight even eighteen. The paper points out that at those levels, the theoretical guarantee basically falls apart — you can’t translate the math into actual privacy protection. So they say you still need to empirically test the synthetic data, even if it was generated with differential privacy.

Meng: That makes sense from an engineering standpoint. If I’m building a system and someone tells me “this is differentially private with epsilon equals ten” I have no idea if that’s actually safe. I’d want to run attacks against it, not just trust the number.

Jane: Exactly, Meng. And that’s why the panel recommends using membership and attribute disclosure metrics instead. Membership disclosure is about whether an attacker can figure out if a specific person was in the training data. Attribute disclosure is about whether an attacker can predict a sensitive attribute, like a diagnosis, for a specific person.

Tom: And the paper even gives practical guidance on how to compute these. For membership disclosure, they say you need to account for the sampling fraction — how big your training set is relative to the whole population. And you should report the F1 score relative to a naive guess, so you know how much better the attacker is doing than just guessing.

Lu: Right, and they even suggest a threshold for that relative score — zero point two — but the panel was actually uncertain about that number. It’s a starting point, not a hard rule.

Meng: So it’s like a baseline for engineers to calibrate against, but you still have to think about your specific data and context.

Jane: Exactly. And for attribute disclosure, they say you need to compare predictions on people who were in the training data versus people who weren’t. That difference tells you whether the synthetic data is leaking individual information or just reflecting general population trends.

Tom: And that distinction is so important — it’s the difference between learning something about a specific person versus learning something about a group. The panel says only the first one counts as a privacy violation.

Lu: That’s a really clean way to separate knowledge generation from disclosure. Research is supposed to generate knowledge, but it shouldn’t reveal individuals.

Jane: Right. And that’s the heart of the framework. Now, let’s talk about what the paper says we should actually improve.

Improvements: Tom: Alright, Jane, we’re back with “A Consensus Privacy Metrics Framework for Synthetic Data,” and this is where the paper really gets practical. The panel doesn’t just say “stop doing this” — they give concrete recommendations for what to do instead.

Jane: Right, and one of the biggest improvements is about quasi-identifiers. Those are the attributes an attacker might know about a person, like age, gender, or zip code. The panel says you should base your privacy metrics on those, not on every single attribute in the dataset.

Lu: And that’s a really interesting point, because a lot of people assume that using all attributes is the worst-case scenario. But the paper shows that’s not true for synthetic data. If you match on too many attributes, you’re more likely to get a mismatch because synthetic data isn’t a perfect copy. So an attacker might actually do better with fewer attributes.

Meng: That’s a counterintuitive result, but it makes sense. If I’m trying to find a match, using more fields means more chances for something to be different. So the adversary would be smart to pick a subset of attributes that gives them the highest success rate.

Jane: Exactly. And that’s why the panel says the data controller needs to decide which attributes are quasi-identifiers based on the context. It’s not a one-size-fits-all thing.

Tom: And then there’s the recommendation about not pre-selecting “vulnerable” records. Some methods only test the risk on records that seem rare or unusual, but the panel says that’s not reliable. You need to evaluate all records because you don’t know which ones are actually at risk.

Lu: Right, and that’s a computational challenge, but it’s the only way to get an honest picture. The paper also suggests reporting vulnerability both for individual synthetic datasets and across multiple datasets from the same model, because synthetic data generation is stochastic — you get different results each time you run it.

Meng: So it’s like running a test suite multiple times to make sure the results are stable, not just a one-off pass.

Jane: Exactly. And then there’s the attribute disclosure improvement — they recommend using a non-member baseline. So you train a model on synthetic data and test it on people who weren’t in the training data. That tells you what the model learns just from population patterns. Then you compare that to how well it predicts for people who were in the training data. The difference is the actual disclosure risk.

Tom: And that’s a really clean way to separate general knowledge from individual leakage. It’s not perfect — the panel admits there’s still work to do on standardizing the prediction models — but it’s a solid step forward.

Lu: And they also flag future research needs, like better identity disclosure metrics and more work on adversarial matching strategies. So this framework is really a starting point, not the final word.

Meng: Which is exactly what we need — a common language to talk about privacy so we can actually compare tools and make better decisions.

Jane: Couldn’t agree more. Let’s wrap this up.

Conclusion: Tom: We’ve been talking about “A Consensus Privacy Metrics Framework for Synthetic Data,” and I think we can all agree this paper is a big deal for anyone working with synthetic data.

Jane: Absolutely. The panel brought together experts from all over the world and gave us a clear direction: stop relying on similarity metrics and epsilon, and focus on membership and attribute disclosure instead.

Lu: And they gave us practical guidance on how to do that — using quasi-identifiers, accounting for sampling fractions, and comparing against non-member baselines. It’s not perfect, but it’s a huge step toward standardization.

Meng: From my side, this is really useful. It means I can actually build privacy evaluations that are comparable across different synthetic data tools, and I can explain to stakeholders why we’re measuring what we’re measuring.

Jane: And that’s the key — this framework gives us a common language. It won’t solve every problem, but it gives us a starting point for making synthetic data sharing safer and more trustworthy.

Tom: And that’s what we need if synthetic data is going to become a mainstream way to share sensitive information, especially in healthcare and government.

Lu: Exactly. The paper also points out where we still need research — identity disclosure metrics, better attack models, and more work on thresholds. So this is really the beginning of a conversation, not the end.

Jane: Well said, Lu. So, thank you to the authors and the panel for putting this together. We’re going to say goodbye to “A Consensus Privacy Metrics Framework for Synthetic Data” and get ready to dive into the next paper.

Tom: Thanks for listening, everyone. See you next time.

More episodes

← Home