Differential Privacy Meets Invariant Statistics: Some Conundrums in Quantifying Trade-Offs
cs.CR, stat.OT
Submitted: 2025-04-21
Updated: 2026-08-31
Comments: 20 pages plus references, 1 figure
Journal ref: Hotz, V. J., Gong, R., & Schmutte, I. M. (Eds.). (2026). Data privacy protection and the conduct of applied research: Methods, approaches, and new findings. University of Chicago Press
License: http://creativecommons.org/licenses/by/4.0/
The gist: This work was inspired by the question of whether data swapping, a popular form of statistical disclosure control used to protect many data products including three recent US Decennial Censuses, can
Terminology
Abstract
This work was inspired by the question of whether data swapping, a popular form of statistical disclosure control used to protect many data products including three recent US Decennial Censuses, can satisfy differential privacy (DP). Given the existence of more than 200 formulations of DP (and counting), as a precondition to answering this question one must precisely specify what it actually means to be DP. Motivated by this observation, we first conduct a theoretical investigation into DP's fundamental essence, resulting in a five-building-block system explicating the who, where, what, how and how much aspects of DP. Instantiating this system in the context of the US Decennial Census, we then demonstrate the broad applicability and relevance of DP by comparing a swapping strategy like that used in 2010 with the TopDown Algorithm--the main DP method adopted in the 2020 Census. This chapter provides nontechnical summaries of these two pieces of work (developed elsewhere), as well as extended discussions on a number of issues they unearth that complicate the formulation and the navigation of the so-called privacy-utility trade-off: How can greater awareness of the five building blocks thwart privacy theatrics? How can invariants (statistics that are released as is, without any privacy protection) align with DP's philosophy of relative privacy? How do our results bridging traditional statistical disclosure control and DP allow a data custodian to reap the benefits of both these fields? And how can removing the implicit reliance on aleatoric uncertainty lead to new generalizations of DP? Our ultimate goal with these discussions is to deepen the theoretical basis, broaden the practical applicability, and reduce the misperception of DP--all without shaking its core foundations.
Sources
- A Refreshment Stirred, Not Shaken: Invariant-Preserving Deployments of Differential Privacy for the U.S. Decennial Census
- The Complexities of Differential Privacy for Survey Data
- A Note on the Misinterpretation of the US Census Re-identification Attack
- Programming Frameworks for Differential Privacy
- Reconstruction Attacks on Aggressive Relaxations of Differential Privacy
- Privately Answering Queries on Skewed Data via Per Record Differential Privacy
- Composition of Differential Privacy & Privacy Amplification by Subsampling
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs