A Consensus Privacy Metrics Framework for Synthetic Data

arXiv:2503.04980 · cs.CR, cs.AI · Submitted 2025-03-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Consensus Privacy Metrics Framework for Synthetic Data".

Jane: The paper was written by Lisa Pilgram, Fida K. Dankar, Jorg Drechsler, Mark Elliot, Josep Domingo-Ferrer et al. from University of Ottawa and CHEO Research Institute and Charité – Universitätsmedizin Berlin and Institute for Employment Research and Ludwig-Maximilians-Universität and University of Maryland and University of Manchester and Universitat Rovira i Virgili and Max Planck Institute for Software Systems and Virginia Tech and University of Alberta and Vanderbilt University Medical Center and Vanderbilt University and University of Oklahoma and Medicines and Healthcare products Regulatory Agency and Berlin Institute of Health at Charité – Universitätsmedizin Berlin and University Hospital Lausanne.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, we’ve got a paper today that’s going to make a lot of people in data science sit up straight. It’s called “A Consensus Privacy Metrics Framework for Synthetic Data,” and honestly, the title alone tells you we’re finally getting serious about how we measure privacy.

Jane: Absolutely, Tom. And I love that this isn’t just one lab’s opinion. You’ve got researchers from universities, government agencies, and even regulators across Canada, Germany, the UK, Spain, the US, Switzerland — a real global panel. They ran a Delphi process, which is basically a structured way to get experts to agree on something through multiple rounds of voting.

Tom: Right, and that’s huge because synthetic data is being shared everywhere now — health records, census data, clinical trials. But until now, there wasn’t a standard way to say, “Yes, this synthetic dataset is safe to release.”

Jane: Exactly. And the paper’s main message is pretty bold: a lot of the metrics people currently use to claim privacy are either misleading or just wrong. They specifically call out similarity metrics — you know, how close a synthetic record is to a real one — and say those shouldn’t be used as standalone privacy guarantees.

Tom: That’s a big deal because those are the most common metrics in the literature. I mean, if you’re a company selling synthetic data and you’re saying, “Look, our data is private because no synthetic row exactly matches a real row,” this paper is basically saying that’s not enough.

Jane: Right, because similarity doesn’t capture what actually matters — whether someone can figure out if a person was in the training data, or infer something sensitive about them. That’s what they call membership disclosure and attribute disclosure.

Tom: And those are the two types of disclosure the panel says we should focus on. They’re not saying synthetic data is unsafe, just that we need better tools to measure it.

Jane: And that’s what the rest of the paper is about — giving us those tools, or at least the framework to build them. I’m excited to dig into the details.

Tom: Me too. Let’s get into the summary and see what the panel actually recommends.

Summary: Jane: So, Tom, we’re back with “A Consensus Privacy Metrics Framework for Synthetic Data,” and the summary really lays out the core findings. The panel basically says we need to stop relying on similarity metrics and the privacy budget epsilon from differential privacy as the main ways to report privacy.

Tom: Yeah, and that epsilon point is spicy. For people who don’t know, differential privacy gives you a number, epsilon, that’s supposed to tell you how much privacy you’re getting. Lower is better. But the panel says unless epsilon is really close to zero, that number doesn’t mean much in practice.

Lu: And that’s a really important nuance, Tom. A lot of real-world systems use epsilon values like four eight even eighteen. The paper points out that at those levels, the theoretical guarantee basically falls apart — you can’t translate the math into actual privacy protection. So they say you still need to empirically test the synthetic data, even if it was generated with differential privacy.

Meng: That makes sense from an engineering standpoint. If I’m building a system and someone tells me “this is differentially private with epsilon equals ten” I have no idea if that’s actually safe. I’d want to run attacks against it, not just trust the number.

Jane: Exactly, Meng. And that’s why the panel recommends using membership and attribute disclosure metrics instead. Membership disclosure is about whether an attacker can figure out if a specific person was in the training data. Attribute disclosure is about whether an attacker can predict a sensitive attribute, like a diagnosis, for a specific person.

Tom: And the paper even gives practical guidance on how to compute these. For membership disclosure, they say you need to account for the sampling fraction — how big your training set is relative to the whole population. And you should report the F1 score relative to a naive guess, so you know how much better the attacker is doing than just guessing.

Lu: Right, and they even suggest a threshold for that relative score — zero point two — but the panel was actually uncertain about that number. It’s a starting point, not a hard rule.

Meng: So it’s like a baseline for engineers to calibrate against, but you still have to think about your specific data and context.

Jane: Exactly. And for attribute disclosure, they say you need to compare predictions on people who were in the training data versus people who weren’t. That difference tells you whether the synthetic data is leaking individual information or just reflecting general population trends.

Tom: And that distinction is so important — it’s the difference between learning something about a specific person versus learning something about a group. The panel says only the first one counts as a privacy violation.

Lu: That’s a really clean way to separate knowledge generation from disclosure. Research is supposed to generate knowledge, but it shouldn’t reveal individuals.

Jane: Right. And that’s the heart of the framework. Now, let’s talk about what the paper says we should actually improve.

Improvements: Tom: Alright, Jane, we’re back with “A Consensus Privacy Metrics Framework for Synthetic Data,” and this is where the paper really gets practical. The panel doesn’t just say “stop doing this” — they give concrete recommendations for what to do instead.

Jane: Right, and one of the biggest improvements is about quasi-identifiers. Those are the attributes an attacker might know about a person, like age, gender, or zip code. The panel says you should base your privacy metrics on those, not on every single attribute in the dataset.

Lu: And that’s a really interesting point, because a lot of people assume that using all attributes is the worst-case scenario. But the paper shows that’s not true for synthetic data. If you match on too many attributes, you’re more likely to get a mismatch because synthetic data isn’t a perfect copy. So an attacker might actually do better with fewer attributes.

Meng: That’s a counterintuitive result, but it makes sense. If I’m trying to find a match, using more fields means more chances for something to be different. So the adversary would be smart to pick a subset of attributes that gives them the highest success rate.

Jane: Exactly. And that’s why the panel says the data controller needs to decide which attributes are quasi-identifiers based on the context. It’s not a one-size-fits-all thing.

Tom: And then there’s the recommendation about not pre-selecting “vulnerable” records. Some methods only test the risk on records that seem rare or unusual, but the panel says that’s not reliable. You need to evaluate all records because you don’t know which ones are actually at risk.

Lu: Right, and that’s a computational challenge, but it’s the only way to get an honest picture. The paper also suggests reporting vulnerability both for individual synthetic datasets and across multiple datasets from the same model, because synthetic data generation is stochastic — you get different results each time you run it.

Meng: So it’s like running a test suite multiple times to make sure the results are stable, not just a one-off pass.

Jane: Exactly. And then there’s the attribute disclosure improvement — they recommend using a non-member baseline. So you train a model on synthetic data and test it on people who weren’t in the training data. That tells you what the model learns just from population patterns. Then you compare that to how well it predicts for people who were in the training data. The difference is the actual disclosure risk.

Tom: And that’s a really clean way to separate general knowledge from individual leakage. It’s not perfect — the panel admits there’s still work to do on standardizing the prediction models — but it’s a solid step forward.

Lu: And they also flag future research needs, like better identity disclosure metrics and more work on adversarial matching strategies. So this framework is really a starting point, not the final word.

Meng: Which is exactly what we need — a common language to talk about privacy so we can actually compare tools and make better decisions.

Jane: Couldn’t agree more. Let’s wrap this up.

Conclusion: Tom: We’ve been talking about “A Consensus Privacy Metrics Framework for Synthetic Data,” and I think we can all agree this paper is a big deal for anyone working with synthetic data.

Jane: Absolutely. The panel brought together experts from all over the world and gave us a clear direction: stop relying on similarity metrics and epsilon, and focus on membership and attribute disclosure instead.

Lu: And they gave us practical guidance on how to do that — using quasi-identifiers, accounting for sampling fractions, and comparing against non-member baselines. It’s not perfect, but it’s a huge step toward standardization.

Meng: From my side, this is really useful. It means I can actually build privacy evaluations that are comparable across different synthetic data tools, and I can explain to stakeholders why we’re measuring what we’re measuring.

Jane: And that’s the key — this framework gives us a common language. It won’t solve every problem, but it gives us a starting point for making synthetic data sharing safer and more trustworthy.

Tom: And that’s what we need if synthetic data is going to become a mainstream way to share sensitive information, especially in healthcare and government.

Lu: Exactly. The paper also points out where we still need research — identity disclosure metrics, better attack models, and more work on thresholds. So this is really the beginning of a conversation, not the end.

Jane: Well said, Lu. So, thank you to the authors and the panel for putting this together. We’re going to say goodbye to “A Consensus Privacy Metrics Framework for Synthetic Data” and get ready to dive into the next paper.

Tom: Thanks for listening, everyone. See you next time.

Lisa Pilgram, Fida K. Dankar, Jorg Drechsler, Mark Elliot, Josep Domingo-Ferrer, Paul Francis, Murat Kantarcioglu, Linglong Kong, Bradley Malin, Krishnamurty Muralidhar, Puja Myles, Fabian Prasser, Jean Louis Raisaro, Chao Yan, Khaled El Emam

University of Ottawa · CHEO Research Institute · Charité – Universitätsmedizin Berlin · Institute for Employment Research · Ludwig-Maximilians-Universität · University of Maryland · University of Manchester · Universitat Rovira i Virgili · Max Planck Institute for Software Systems · Virginia Tech · University of Alberta · Vanderbilt University Medical Center · Vanderbilt University · University of Oklahoma · Medicines and Healthcare products Regulatory Agency · Berlin Institute of Health at Charité – Universitätsmedizin Berlin · University Hospital Lausanne

cs.CR, cs.AI

Submitted: 2025-03-06

Updated: 2026-08-18

Journal ref: Patterns, Volume 6, Issue 10, 101320, 2025

DOI: 10.1016/j.patter.2025.101320

Project page: https://datasciencecampus.github.io/synthgauge/autoapi/synthgauge/metrics/privacy/index.html

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: Synthetic data generation is one approach for sharing individual-level data.

Key concepts

Membership Disclosure
This measures whether an attacker can figure out if a specific person was part of the original training data used to generate the synthetic dataset. The framework suggests calculating this by comparing the F1 score against a naive guess, providing a quantifiable measure of leakage.
Attribute Disclosure
This measures whether an attacker can predict sensitive information, like a diagnosis, for an individual. It is determined by comparing how well the model predicts outcomes for people who were in the training set versus those who were not in the training set.
Differential Privacy (Epsilon)
A mathematical value used to quantify privacy protection. The panel notes that if epsilon is high, such as 4 or 18, the theoretical guarantee does not translate into actual practical privacy protection, requiring empirical testing of the synthetic data.
Quasi-Identifiers
These are attributes an attacker might know about a person, such as age or zip code. The framework recommends basing privacy metrics on these specific attributes rather than using every single attribute in the dataset, as this provides more relevant context for risk assessment.

Terminology

Summary

Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals’ privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.

Improvements for AI systems

Based on the consensus framework and critical analysis in this paper, I can implement the following specific improvements to AI systems, particularly those involved in synthetic data generation (SDG) and privacy evaluation:

1. Implement a Two-Tier Privacy Evaluation System (Membership + Attribute Disclosure)

  • What to build: A mandatory post-processing module for any SDG model (GANs, VAEs, diffusion models, Bayesian networks) that computes both membership disclosure (using the F rel metric) and attribute disclosure (using the non-member baseline) before data release.

  • Specific implementation: The system will automatically split the original data into training/holdout sets, generate multiple synthetic datasets (≥10 to account for stochasticity), and compute:

  • Membership: F1 score, F naive (based on sampling fraction), and F rel = (F1 - F naive)/(1 - F naive). Flag if F rel > 0.2.

  • Attribute: Train a prediction model (e.g., random forest) on synthetic data to predict each sensitive attribute; compute AUROC for members (training set) and non-members (holdout set). Flag if (AUROC member - AUROC nonmember) > 0.15 AND AUROC member > 0.6.

  • Outcome: The AI system can now provide a binary release/no-release recommendation with clear, interpretable thresholds, rather than relying on vague similarity scores.

2. Replace Similarity-Based Privacy Metrics with Disclosure-Specific Metrics

  • What to fix: Current systems often report distance to closest record (DCR) or exact match count as privacy guarantees. The paper proves these are misleading (they measure reconstruction, not identity disclosure).

  • Specific implementation: Remove all record-level similarity metrics (e.g., Euclidean distance, Hamming distance) from the privacy evaluation pipeline. Instead, the system will only use the two-tier evaluation above. If a user requests a similarity metric, the system will return an error explaining why it's not a valid privacy indicator.

  • Outcome: Prevents false confidence in synthetic data that is similar but actually leaks sensitive attributes.

3. Implement Quasi-Identifier (QI) Selection and Subset Search for Worst-Case Vulnerability

  • What to build: A module that, instead of using all attributes (which the paper shows underestimates risk), identifies QIs based on replicability, distinguishability, and knowability criteria.

  • Specific implementation: The system will:

  • Accept a user-defined list of QIs or auto-detect them using heuristics (e.g., low cardinality, high stability over time).

  • For membership disclosure, perform an exhaustive search over all subsets of QIs (up to a computational limit, e.g., 15 attributes) and all generalization levels (e.g., age bins, geographic region rollups) to find the maximum F rel.

  • Report both the average and the maximum vulnerability to avoid underestimation.

  • Outcome: The AI system provides a true worst-case privacy estimate, preventing an adversary from exploiting a subset of attributes that the controller missed.

4. Correct the Attack Dataset Composition for Membership Disclosure

  • What to fix: Most current implementations use a 50/50 split of members/non-members in the attack dataset, which is unrealistic. The paper shows this overestimates risk when the sampling fraction is low.

  • Specific implementation: The system will automatically set the member prevalence in the attack dataset equal to the actual sampling fraction (n/N) of the training data from the population. It will also compute and report F naive (the adversary's success by random guessing) so the user can see the incremental risk added by the synthetic data.

  • Outcome: More accurate, context-aware vulnerability estimates. For example, if a dataset represents 5% of the population, the system will not falsely report a high F1 score that would occur with a 50/50 split.

5. Integrate a Differential Privacy (DP) Audit Module

  • What to build: A validation layer for DP-based SDG models. The paper shows that epsilon values > 0 are not interpretable and that even small epsilon values can have implementation flaws.

  • Specific implementation: The system will:

  • Reject any claim of privacy based solely on epsilon (unless epsilon < 0.1, and even then, require empirical validation).

  • For any DP synthetic data, run the same membership and attribute disclosure evaluations as for non-DP data.

  • Flag if the empirical vulnerability exceeds the theoretical guarantee (e.g., if F rel > 0.2 despite a claimed epsilon of 0.1).

  • Outcome: Prevents privacy theater where a system claims DP protection but actually leaks information due to implementation errors or overly large epsilon.

6. Implement a Non-Member Baseline for Attribute Disclosure (Knowledge Generation vs. Privacy)

  • What to build: A module that separates legitimate knowledge generation (e.g., learning that smokers have higher cancer risk) from privacy violations (learning an individual's specific diagnosis).

  • Specific implementation: For each sensitive attribute, the system will:

  • Train a prediction model on synthetic data.

  • Evaluate AUROC on (a) members (training set) and (b) non-members (holdout set).

  • Report the difference (A rel = AUROC member - AUROC nonmember) and the absolute member AUROC.

  • Only flag as a privacy violation if A rel > 0.15 AND AUROC member > 0.6. If the model predicts well for both members and non-members, it's classified as knowledge generation, not disclosure.

  • Outcome: Allows the AI system to share valuable synthetic data that preserves population-level insights without penalizing it for being too accurate on unseen data.

7. Add a Stochasticity-Aware Reporting Module

  • What to build: A wrapper that generates multiple synthetic datasets (default = 10) from the same trained model and reports the mean and standard deviation of all privacy metrics.

  • Specific implementation: The system will:

  • Run the SDG model 10 times with different random seeds.

  • Compute membership and attribute disclosure metrics for each run.

  • Report the average and 95% confidence interval, so users know if the vulnerability is stable or varies wildly.

  • Outcome: Prevents a single lucky/unlucky synthetic dataset from skewing the privacy assessment, leading to more reliable release decisions.

8. Implement a Context-Aware Threshold Adjustment Engine

  • What to build: A decision-support tool that adjusts the default thresholds (F rel ≤ 0.2, A rel ≤ 0.15, AUROC member ≤ 0.6) based on the Invasion of Privacy construct.

  • Specific implementation: The system will accept inputs on:

  • Data sensitivity (e.g., HIV status vs. favorite color).

  • Potential harm (e.g., legal prosecution vs. embarrassment).

  • Consent/notice (e.g., explicit opt-in vs. secondary use without notice).

  • Benefit to society (e.g., pandemic response vs. marketing).

  • It will then recommend a more conservative (e.g., F rel ≤ 0.1) or permissive (e.g., F rel ≤ 0.3) threshold, with a clear justification.

  • Outcome: Makes the privacy framework adaptable to real-world scenarios, avoiding a one-size-fits-all approach that could be too strict (blocking beneficial data) or too lenient (allowing harmful disclosures).

9. Build a Vulnerable Record Search Algorithm

  • What to fix: The paper shows that pre-selecting vulnerable records (e.g., rare outliers) is flawed. The system should not assume which records are at risk.

  • Specific implementation: The system will compute disclosure vulnerability for all records in the dataset (not a subset) and then identify the top 1% with the highest vulnerability. This is computationally intensive but necessary for a correct worst-case assessment. The system will use efficient nearest-neighbor search (e.g., KD-trees) to make this feasible for datasets up to 1 million rows.

  • Outcome: Ensures that no high-risk record is missed due to incorrect assumptions about what makes a record vulnerable.

10. Provide a Standardized Reporting Format for Regulators

  • What to build: A structured output (e.g., JSON or PDF) that includes all recommended metrics, thresholds, and context adjustments, making it easy for data controllers to demonstrate compliance with privacy regulations.

  • Specific implementation: The report will include:

  • QI list and justification.

  • Number of synthetic datasets evaluated.

  • F naive, F1, F rel (with max and average).

  • AUROC member, AUROC nonmember, A rel (for each sensitive attribute).

  • DP epsilon (if applicable) and empirical validation results.

  • Final release recommendation (release/no-release) with rationale.

  • Outcome: Reduces regulatory uncertainty and speeds up the approval process for sharing synthetic health data.

Summary of What the Improved AI System Can Do:

  • Accurately assess privacy without relying on misleading similarity metrics.

  • Provide worst-case vulnerability estimates by searching over QI subsets and generalizations.

  • Distinguish knowledge generation from privacy violations, allowing safe sharing of valuable population-level insights.

  • Correctly model adversary behavior by using realistic attack dataset compositions.

  • Validate DP claims empirically, preventing privacy theater.

  • Adapt thresholds to context, balancing privacy with societal benefit.

  • Produce regulator-ready reports, facilitating faster and safer data sharing for research.

Abstract

Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.

Sources

Related papers