Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Vinícius Gabriel Angelozzi, Héber H. Arcolezi
cs.LG, cs.AI, cs.CR
Submitted: 2026-07-08
Comments: Paper accepted at PETS 2026. Code is available at https://github.com/vinicius-verona/dp-fair-intervention-benchmark
Code: https://github.com/vinicius-verona/dp-fair-intervention-benchmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
Terminology
Abstract
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups. However, these objectives can conflict: DP often amplifies disparities across demographic groups, and little is known about whether established fairness interventions remain effective under DP constraints. In this work, we present, to our knowledge, the first systematic evaluation of fairness interventions on differentially private synthetic tabular data. Our benchmark centers on the Adaptive Iterative Mechanism (AIM), identified as the state-of-the-art marginal-based DP synthesizer (Cormode et al. 2025). We thus evaluate fairness interventions across four datasets, multiple group fairness metrics, and three categories of mitigation strategies (pre-processing, in-processing, and post-processing) under a wide range of privacy budgets. We compare four pipeline configurations: (Baseline) training on original data; (DP-only) training on DP synthetic data; (Fair-only) applying fairness mechanisms on original data; and (DP+Fair) combining fairness mechanisms with DP synthetic data. Our results demonstrate that while DP alone can degrade both utility and fairness, applying fairness interventions can partially restore equitable outcomes. Among them, post-processing methods tend to provide more stable fairness-utility trade-offs across privacy budgets and synthesizers, achieving strong fairness improvements while preserving competitive utility relative to other intervention stages. We release all code, data, and experimental artifacts in an open-source repository to ensure full reproducibility and to support future research on the privacy-fairness-utility trade-off.
Sources
- Synthetic Data -- what, why and how?
- Evaluating the Fairness Impact of Differentially Private Synthetic Data
- Why Does Differential Privacy with Large Epsilon Defend Against Practical Membership Inference Attacks?
- Winning the NIST Contest: A scalable and general approach to differentially private synthetic data
- Benchmarking Differentially Private Synthetic Data Generation Algorithms
- DP-SGD vs PATE: Which Has Less Disparate Impact on Model Accuracy?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks