LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings
cs.LG, cs.CR
Submitted: 2026-02-21
Updated: 2026-09-09
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Membership inference attacks (MIAs) threaten the privacy of machine learning models by revealing whether a data point was used during training.
Terminology
Abstract
Membership inference attacks (MIAs) threaten the privacy of machine learning models by revealing whether a data point was used during training. Existing MIAs often assume access to public datasets, shadow models, confidence scores or the training distribution, which makes them vulnerable to defenses like confidence masking. Label-only MIAs avoid these assumptions but require thousands of queries per sample. We propose a cost-effective label-only MIA framework based on transferability and model extraction. Querying the target M with active sampling, perturbation-based selection and synthetic data, we extract a surrogate S on which membership inference is performed offline. This shifts query overhead to a one-time extraction phase. It also removes the restriction that defines the label-only setting: the attacker controls S and can read its posteriors and training trajectory, so attacks that cannot be run against M can be run against S. On Location, Purchase and Texas, the strongest attack on S improves AUC over the direct attack on M by 0.9, 5.6 and 5.0 percentage points, and improves the true positive rate at 1% false positive rate by 2.8 times to 6.3 times. We characterize how leakage transfer depends on surrogate fidelity, evaluate standard defenses, and report preliminary results on image datasets.
Sources
- Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
- Defending Model Inversion and Membership Inference Attacks via Prediction Purification
- Do Membership Inference Attacks Work on Large Language Models?
- Data-Free Model Extraction
- Adam: A Method for Stochastic Optimization
- Adversarial Robustness Toolbox v1.0.0
- OSLO: One-Shot Label-Only Membership Inference Attacks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- CINIC-10 is not ImageNet or CIFAR-10
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks