QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing
cs.LG, cs.CR
Submitted: 2026-09-16
Updated: 2026-09-18
Code: https://github.com/wsqwsq/QuanText
License: http://creativecommons.org/licenses/by/4.0/
The gist: Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of
Terminology
Abstract
Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of records associated with a particular gender, diagnosis, or political stance. Existing work has largely focused on property inference attacks that recover such global properties, while defenses for protecting these dataset-level secrets remain limited. Differential privacy, although effective for protecting individual records, provides only weak protection for aggregate properties. We propose Randomized Quantization for Text (QuanText), a training-free and large-language-model-agnostic data release mechanism that protects global secrets in textual datasets while preserving data utility. Given a dataset-level secret, such as the proportion of records with a particular diagnosis, and attributes whose utility should be preserved, such as topic and sentiment, QuanText perturbs both the secret distribution and the distributions of correlated attributes. It does so by constructing candidate release distributions over secret and non-secret attributes, randomly selecting a candidate sufficiently close to the private empirical distribution, and rewriting each private text sample to match the selected distribution using attribute-related snippets from the original text. QuanText is inspired by the Statistic Maximal Leakage (SML) framework, which bounds leakage about a secret function of a data distribution. Under idealized conditions, we show that QuanText satisfies an SML guarantee. Since these conditions may not hold exactly in practice, we also evaluate QuanText empirically on real-world datasets. Our results show that QuanText achieves a better empirical privacy-utility trade-off than competing data generation baselines.
Sources
- Property Inference Attacks Against GANs
- Can We Infer Confidential Properties of Training Data from LLMs?
- PriSampler: Mitigating Property Inference of Diffusion Models
- On the (In)Feasibility of Attribute Inference Attacks on Machine Learning Models
- Formalizing and Estimating Distribution Inference Risks
- Lessons Learned: Defending Against Property Inference Attacks
- Differentially Private Synthetic Data via Foundation Model APIs 2: Text
- Differentially Private Synthetic Data via APIs 3: Using Simulators Instead of Foundation Model
- Differentially Private Fine-tuning of Language Models
- DPImageBench: A Unified Benchmark for Differentially Private Image Synthesis
- SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
- Guarding Multiple Secrets: Enhanced Summary Statistic Privacy for Data Sharing
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks