Black-Box Membership Inference via Word-Level Probability Estimation

arXiv:2609.10611 · cs.CR, cs.LG, stat.ML · Submitted 2026-09-08 · Read on arXiv

cs.CR, cs.LG, stat.ML

Submitted: 2026-09-08

Updated: 2026-09-08

Comments: 17 pages, 5 figures, and 17 tables

Code: https://github.com/niusj03/WPMIA

License: http://creativecommons.org/licenses/by/4.0/

The gist: Membership inference attacks (MIAs) have emerged as critical tools for auditing privacy risks in large language models (LLMs), aiming to determine whether a given text was included in a model's

Terminology

Abstract

Membership inference attacks (MIAs) have emerged as critical tools for auditing privacy risks in large language models (LLMs), aiming to determine whether a given text was included in a model's training corpus. However, most existing MIAs require access to per-token logits or probabilities, making them inapplicable in practice to proprietary LLMs that expose only textual continuations. To address this underexplored setting, we propose Word-level Probability MIA (WPMIA), a statistically principled MIA for strict black-box privacy auditing. WPMIA estimates word-level generation probabilities via Monte Carlo sampling with local kernel smoothing, then aggregates these estimates into a sequence-level likelihood estimator. Furthermore, WPMIA constructs the likelihood conditioned on different prefixes, thereby amplifying the distributional differences between members and non-members. We evaluate WPMIA across various open-source LLMs and find that it consistently outperforms existing black-box baselines. Importantly, we also evaluate WPMIA on modern proprietary LLMs, including GPT-5-Chat, Gemini-2.5-Flash, and Claude-4.5-Haiku, achieving an average TPR@5%FPR of 42.0 across these models. These results offer a sound foundation for future research on strict black-box membership inference. Code is available at https://github.com/niusj03/WPMIA.

Related papers