Batch Normalization Amplifies Memorization and Privacy Risks
cs.LG, cs.AI
Submitted: 2026-05-23
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Batch Normalization (BN) is widely adopted to enable faster convergence and more stable training of deep neural networks.
Terminology
Abstract
Batch Normalization (BN) is widely adopted to enable faster convergence and more stable training of deep neural networks. However, its impact on privacy and memorization has remained largely unexplored. In this work, we investigate the effect of BN layers on the memorization of atypical or outlier samples and its implications for privacy leakage. We conduct an extensive empirical study using three complementary approaches: (i) unintended memorization of out-of-distribution samples, (ii) per-sample influence, and (iii) susceptibility to membership inference attacks (MIA). Across multiple datasets and architectures, we consistently observe that BN substantially increases the memorization of outliers compared to models without BN. Critically, this amplified memorization translates directly into privacy vulnerabilities: models with BN exhibit significantly higher susceptibility to MIAs. We complement our empirical findings with a mechanistic analysis under the exact BN backward pass, which shows that BN amplifies the per-step margin growth of outlier samples during training. Our results highlight an underappreciated privacy risk associated with BN and provide both practical and theoretical insights into how normalization layers can amplify the influence of rare or sensitive training examples.
Sources
- Can Neural Network Memorization Be Localized?
- Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
- Quantifying and Localizing Usable Information Leakage from Neural Network Gradients
- On the geometry of generalization and memorization in deep neural networks
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Batch Normalization is a Cause of Adversarial Vulnerability
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Understanding deep learning requires rethinking generalization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks