Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding
cs.LG, cs.CR
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/huyvnphan/PyTorch_CIFAR
License: http://creativecommons.org/licenses/by/4.0/
The gist: Instance encoding is a popular empirical technique for privacy enhancement when sharing data to an untrusted server.
Terminology
Abstract
Instance encoding is a popular empirical technique for privacy enhancement when sharing data to an untrusted server. It transforms sensitive data through an encoding process before sharing, with the hope that the encoding process retains utility but makes it hard to reconstruct the original data. However, most work offers no theoretical guarantee that the encoding process is actually irreversible. A recent work derived a mean-squared error (MSE) bound limiting any adversary's reconstruction accuracy, offering one of the first theoretical results in this domain. This bound, however, has three critical limitations: it is often too loose, only works with randomized encoders (excluding many deterministic encoders practitioners use), and only bounds MSE. We introduce a family of new bounds that (1) are tighter, (2) applicable even to fully deterministic encoders, and (3) can extend beyond MSE to other norm-based similarity metrics, by properly accounting for the encoder's spectral structure. We evaluate our bounds across a range of encoders, datasets, and attacks, showing they hold consistently and improve upon the existing bound.
Sources
- NeuraCrypt is not private
- Inferential Privacy Guarantees for Differentially Private Mechanisms
- Measuring Data Leakage in Machine-Learning Models with Fisher Information
- NICE: Non-linear Independent Components Estimation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Denoising Diffusion Probabilistic Models
- PrivyNet: A Flexible Framework for Privacy-Preserving Deep Neural Network Training
- Enabling Inference Privacy with Adaptive Noise Injection
- Score-Based Generative Modeling through Stochastic Differential Equations
- Practical Defences Against Model Inversion Attacks for Split Neural Networks
- Split Learning for collaborative deep learning in healthcare
- Split learning for health: Distributed deep learning without sharing raw patient data
- NeuraCrypt: Hiding Private Health Data via Random Neural Networks for Public Training
- Syfer: Neural Obfuscation for Private Data Release
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks