Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

arXiv:2608.09140 · cs.CR, cs.CL · Submitted 2026-08-10 · Read on arXiv

Li Siyan, Zhou Yu, Julia Hirschberg

Columbia University

cs.CR, cs.CL

Submitted: 2026-08-10

Updated: 2026-08-11

Comments: Accepted into HAIPS workshop at COLM 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors.

Terminology

Summary

Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors. One such approach is Privacy-Conscious Delegation (PCD), where a local LLM acts as an intermediary. However, privacy risk does not stem solely from explicit identifiers but also PII-free self-disclosures, leaving users identifiable through combinations of quasi-identifying traits. We investigate a probabilistic variant of PCD, where we augment its objectives with an LLM-driven probabilistic estimation of k-anonymity. To facilitate this, we first create the PUPA-SD dataset, which contains naturalistic user queries with self-disclosure. Our preliminary results indicate that optimizing PAPILLON on PUPA-SD improves quality on unseen conversations across a variety of local models and produces the best privacy-utility balance for Llama-3.2-3B, while smaller models struggle to jointly optimize quality and privacy. We propose k-anonymity as a useful auxiliary metric for tackling PCD.

LLM-powered agents and interactive systems pose significant privacy risks, not only induced by memorization of user information from model training and therefore vulnerability to data extraction attacks, but also by self-disclosure during human-LLM interactions. To exacerbate the matter, as LLMs exhibit more and more human-like traits in conversation, such as empathy and matching user linguistic styles, users develop greater trust in the models, increasing their willingness to share intimate personal details. Sharing private information increases inference-time privacy risks, since these self-disclosures could be used for the next round of training or divulged due to server-side data breaches.

Prior work on protecting users from inference-time privacy risks has focused on reducing the exposure of PIIs and explicit identifiers to untrusted, remote frontier models through text sanitization and abstraction or by leveraging trusted, locally hosted models. One framework that formalizes this task structure is Privacy-Conscious Delegation (PCD). PAPILLON serves as a strong baseline for PCD. It is important to note, however, that eliminating explicit identifiers may not fully preserve anonymity: users can remain identifiable via quasi-identifiers (traits whose combinations uniquely narrow the anonymity set). The notion of k-anonymity formalizes this risk in terms of the ease of pinpointing a person given a set of traits with respect to a known population. To operationalize anonymity reasoning for free-form texts, Zheng et al. (2025) introduces BRANCH, a probabilistic framework that leverages LLMs' knowledge of census data and logical dependencies to estimate k-anonymity with relatively high accuracy. BRANCH therefore allows us to integrate k-anonymity into PAPILLON-based pipelines, leveraging the flexibility of prompt optimization.

In this work, we take a first step toward probabilistic privacy-conscious delegation by augmenting PCD objectives beyond explicit identifier leakage. Concretely, we begin with PAPILLON as a strong PCD framework and integrate BRANCH-based probabilistic k-anonymity as an additional objective. In order to optimize and evaluate our probabilistic framework, we extract PUPA-SD dataset from subsets of WildChat and LMSYS-Chat-1M, containing 166 real user queries containing self-disclosures. We show that optimizing for k-anonymity on PUPA-SD can transfer to held-out PUPA-TNB, with Llama-3.2-3B achieving the strongest privacy-utility balance after optimization (Quality: 67.4, PII Leakage: 11.0).

In PCD, given a user query q containing private information units p1, p2,..., pn, local model MLOCAL, and remote model MREMOTE, a PCD system must ensure MREMOTE receives as little private information as possible while maintaining comparable response quality to passing q into MREMOTE directly. PAPILLON provides a strong, prompt-optimization-based baseline for PCD. Given q, MLOCAL produces a redacted query q' that is sent to MREMOTE, whose response is modified by MLOCAL to cater to q. The pipeline can be prompt optimized, maintaining quality on 85.5% of user queries while leaking only 7.5% of information when using Llama-3.1-8B-Instruct as MLOCAL. PAPILLON presents one approach to addressing PCD, but it is not the only one; for instance, instead of firing a query to MREMOTE for every user query, MLOCAL could be more selective. PUPA is created from extracting PII-containing user utterances from WildChat. We use its held-out set, PUPA-TNB, for evaluation.

Estimating k-anonymity over free-form text is difficult because disclosures are unstructured and statistically interdependent. BRANCH addresses this by modeling a text as a joint distribution over personal attributes. Given a text, it first selects the disclosures that can plausibly be estimated from population statistics (e.g., nurse, Austin, TX, pregnant), then elicits a Bayesian network over them: an LLM chooses an ordering of the disclosures as random variables and, for each one, determines which of the variables it is conditionally dependent on. This factors the joint distribution into probability terms, each of which is converted into a natural-language query and estimated by an LLM from its pretraining knowledge of demographic statistics. Recombining these estimates and multiplying by a population size yields the expected number of people matching the disclosed attributes. Because the estimator consumes only text and returns a scalar, it can be directly incorporated into a prompt-optimization objective.

Following PUPA, the dataset used to construct and evaluate PAPILLON for PCD, we construct PUPA-Self-Disclosure (PUPA-SD) from initial conversation turns in WildChat and LMSYS-Chat-1M. For feasible manual inspection, we limit each dataset to its first 20,000 instances. To quickly identify conversations with self-disclosure and improve the evaluation pipeline's accuracy, we develop a validated self-disclosure extraction module. We choose to create our own module rather than using existing disclosure extraction models because they are often trained on Reddit-style posts, which may represent a significant shift in distribution from conversation data. Because the extraction process handles large data volumes, we need a fast model for quick evaluation and improvement. Therefore, we test three models with generally high performance and decent speed: GPT-4.1-mini, GPT-5-mini, and GPT-4o-mini. We define seed prompts for a self-disclosure extractor and evaluate each model on 50 synthetic texts with human-annotated disclosure from Zheng et al. (2025) to identify the most suitable model. An GPT-5-based LLM judge is fixed to compare alternatives using the F-1 score to identify disclosures. Results are documented in Table 1: GPT-4.1-mini (Prec. 77.3, Rec. 69.9, F-1 70.3), GPT-5-mini (Prec. 66.6, Rec. 79.3, F-1 70.7), GPT-4o-mini (Prec. 74.6, Rec. 51.7, F-1 58.9). Given its superior performance, we select GPT-5-mini for all disclosure extraction and ordering tasks. We further optimize this module using DSPy's MIPROv2 and GPT-5 to generate 10 candidate prompt sets, selecting the best-performing combination (Disclosure F-1 = 72.5).

We apply the finalized extraction pipeline to the 40,000 initial conversation turns from WildChat and LMSYS-Chat-1M, and additionally filter out non-English and sexual instances during manual inspection. This process yielded a set of 166 turns. We aim to perform further processing and extraction in the future to grow the dataset size. PUPA-SD contains PII-free instances. To quantify this, we use GLiNER-PII, a state-of-the-art PII extraction model from NVIDIA, to extract information, including emails, names, company names, and locations, from PUPA-SD queries. We present the distribution of the number of extracted PIIs in Figure 2. This further emphasizes that eliminating PII leakage does not fully prevent self-disclosure.

Similar to the original definition of Privacy-Conscious Delegation, our modified task involves utilizing both a trusted but weaker model, MLOCAL, and an untrusted but stronger one, MREMOTE. MLOCAL produces a query q' based on the private user query q containing PII units p1, p2,...pn. Here, in addition to preserving the quality of the final response with respect to the original response from the LLM, as well as reducing the amount of PII unit leakage to MREMOTE, we additionally aim to increase k-anonymity of information passed to MREMOTE. To evaluate our task, we employ the same validated LLM judges from Siyan et al. (2025a); additionally, we implement a version of BRANCH to estimate k-anonymity.

BRANCH is a heavily LLM-dependent approach for estimating privacy risk. While accurate, it is costly, motivating improved data efficiency for swift evaluation of our probabilistic PAPILLON pipelines. We adopt BRANCH's disclosure ordering and query generation stages, simplifying the elicited Bayesian structure to a sequence of cumulative conditioning groups, in which disclosures entering the same group are treated as conditionally independent given the preceding group. We apply a minimally altered version of the self-disclosure extractor, modified to extract disclosures about any individual. We then apply the LLM-based disclosure ordering and query generation from Zheng et al. (2025). GPT-5-mini estimates either the total originating population or the percentage of individuals with given attributes; if insufficient information exists, a fallback population of 400M (English speakers) is used. To reduce redundant LLM calls, queries and normalized answers are indexed using openai/text-embedding-3-small: a new query with cosine similarity ≥ 0.95 to a stored query reuses its cached result. k-anonymity is then computed by sequentially multiplying the population estimate by each percentage, as in BRANCH.

In addition to the final response quality, PII leakage, and prompt well-formedness metrics from Siyan et al. (2025a), we instantiate a k-anonymity metric. To scale the k-anonymity metric to the [0,1] range in accordance with PAPILLON, we note that while dividing by the total population (400M) is a natural normalization, it yields near-zero values where meaningfully different anonymity levels (e.g., k = 1,000 vs. k = 100,000) become indistinguishable. We therefore define the reward as log2(k)/log2(400M), which compresses the range while preserving meaningful distinctions across anonymity levels.

To ensure our probabilistic k-anonymity estimator behaves as expected, we examine whether estimated k-anonymity negatively correlates with the degree of self-disclosure in the original query. Intuitively, queries involving more PIIs should narrow the anonymity set and yield lower k-anonymity estimates. We compute the Spearman correlation between the number of PIIs per query and the estimated k-anonymity on PUPA-TNB, which contains a higher density of PII and self-disclosure than PUPA-SD, obtaining ρ = −0.4045 (p < 10−12). We emphasize that this correlation is expected by construction: because the estimate is a function of the detected disclosures, a negative relationship confirms the pipeline's monotonic behavior but does not establish that the estimations are calibrated against real statistics. We therefore treat this as a sanity check rather than formal validation, and discuss calibration as a limitation.

We adopt PAPILLON as our PCD-k pipeline unchanged, augmenting only the optimization objective to additionally account for k-anonymity. For prompt optimization, we apply DSPy's SIMBA prompt optimizer. Given a metric and a threshold of success, SIMBA samples minibatches of data, identifies difficult examples via metric variance, and then either updates the prompt via self-reflection or adds successful input-output pairs as in-context demonstrations. We limit the number of demonstrations to one. Throughout our prompt optimization experiments, the metric is computed as follows: MetricPAPILLON + k Anon(q'), where MetricPAPILLON is a weighted sum of response quality and prompt well-formedness, i.e., the utility term. PII leakage is not included during optimization, but we measure it during evaluation. This composite metric value is normalized. We optimize on PUPA-SD and evaluate on PUPA-TNB, testing Llama-3.2- 1B,3B-Instruct, Llama-3.1-8B-Instruct, and Qwen2.5- 0.5B,1.5B,7B-Instruct as MLOCAL, with GPT-4o-mini fixed as MREMOTE and GPT-5-mini as the LLM judge.

Figure 3 and Table 2 report results across all local models before and after SIMBA optimization on PUPA-SD, evaluated on the held-out PUPA-TNB. Optimization generally improves response quality, with the largest gains observed in Llama-3.2-3B (+13.2) and Qwen-2.5-1.5B (+17.3). PII leakage results are mixed: Llama-3.2-3B achieves a substantial reduction (29.1 → 11.0), while Qwen models show increased leakage after optimization, suggesting that smaller Qwen models struggle to jointly optimize quality and privacy. This trade-off is similarly observed for Llama-3.1-8B-Instruct, where the significant improvement in quality is accompanied by increased PII leakage. Qwen-2.5-7B is the only model where quality degrades after optimization (83.4 → 79.7), potentially due to the model already performing near its ceiling before optimization. k-anonymity scores remain largely stable across all models and conditions, which could be related to rewritten user queries either no longer containing disclosures or becoming impossible to estimate using publicly available data, resulting in the estimator falling back to the full population. One possible explanation of the heterogeneity of improvements across models is that, with the k-anonymity term contributing little variation and PII leakage excluded from the metric by construction, the objective is driven largely by the quality metric. We note that Llama-3.2-1B does receive a lower k-anonymity score (72.5), indicating that the metric remains discriminative when disclosures survive the rewrite, so this account is unlikely to apply uniformly. We leave a direct measurement of estimator fallback rates to future work.

In this work, we take a first step toward probabilistic privacy-conscious delegation by augmenting PCD beyond PII leakage to incorporate k-anonymity as an objective. To support this, we introduce PUPA-SD and implement an efficient BRANCH-based k-anonymity estimator. Our experiments show that optimization on PUPA-SD yields improvements in quality for most models, but not all models successfully reduce leakage, suggesting that model capacity and benchmark headroom are key factors in navigating the privacy-utility tradeoff. Future work may examine the effect of optimizing on both k-anonymity and PII leakage metrics. We hope this work motivates future research into probabilistic privacy risk estimation for human-LLM interaction, particularly as users increasingly share sensitive personal context with frontier models they do not control.

One core limitation is the potential inaccuracy in various parts of the evaluation pipeline. To begin with, BRANCH can often under-estimate k-anonymity as a result of the independence assumption. However, this is justifiable as the estimate would be relatively conservative, suitable for privacy preservation purposes. A more problematic issue is that our estimators are not calibrated against or grounded in real population statistics, as the estimation hinges entirely upon LLM-encoded prior knowledge. Due to financial constraints, we could not leverage the most capable models for self-disclosure extraction and k-anonymity estimation, meaning LLM-encoded census knowledge may be outdated or imprecise; future work could incorporate agentic web search to ground estimates in current data. PUPA-SD currently contains only 166 instances, limiting statistical power and precluding post-training approaches that could further boost smaller models' instruction-following ability; we plan to scale the dataset by broadening the extraction pipeline to additional conversations from both existing corpora and new sources. Finally, PUPA-SD and our evaluation pipeline are restricted to English; extending to multilingual settings would broaden applicability, though manual inspection would become increasingly challenging as more languages are incorporated.

Improvements for AI systems

Improvements to AI Systems:

  1. Privacy-Preserving Query Rewriting with k-Anonymity Awareness
  • Integrate a BRANCH-based probabilistic k-anonymity estimator into the local LLM’s prompt-optimization loop (as done in PAPILLON+).

  • The improved system can rewrite user queries to remove quasi-identifying trait combinations (e.g., occupation + location + pregnancy status) that could uniquely identify users, even when no explicit PII (names, emails) is present.

  • It can dynamically balance response quality and anonymity by optimizing a composite metric (quality + log-scaled k-anonymity), ensuring the remote model receives queries that are both useful and less linkable to a specific individual.

  1. Self-Disclosure Extraction and Risk Scoring for Conversational AI
  • Deploy a fine-tuned GPT-5-mini-based disclosure extractor (optimized via DSPy MIPROv2) that identifies PII-free self-disclosures (e.g., health conditions, family details, lifestyle habits) in real-time user-LLM dialogues.

  • The improved system can flag high-risk disclosures and automatically trigger query sanitization or refuse to forward sensitive context to untrusted remote models, reducing inference-time privacy leakage.

  1. Transferable Privacy-Utility Optimization Across Model Sizes
  • Use PUPA-SD as a training/evaluation benchmark to prompt-optimize local models of varying capacities (0.5B–8B) for privacy-utility trade-offs.

  • The improved system can automatically select the best local model for a given privacy budget (e.g., Llama-3.2-3B achieves quality 67.4 with PII leakage 11.0), enabling deployment of smaller, faster models that still meet privacy standards.

  1. Efficient k-Anonymity Estimation with Caching and Fallback Logic
  • Implement the simplified BRANCH estimator with embedding-based query caching (cosine similarity ≥ 0.95) and a fallback population of 400M English speakers when data is insufficient.

  • The improved system can estimate anonymity sets for free-form text in near-real-time, enabling interactive privacy checks during live conversations without incurring high LLM call costs.

  1. Privacy-Aware Prompt Optimization with Demonstration Selection
  • Use SIMBA optimizer with a composite metric (quality + k-anonymity) and one in-context demonstration to iteratively refine the local model’s rewriting instructions.

  • The improved system can learn from difficult examples (high metric variance) to produce rewrites that preserve utility while maximizing anonymity, even for unseen query distributions (as validated on PUPA-TNB).

  1. Calibrated Privacy Risk Monitoring via Correlation Sanity Checks
  • Integrate a Spearman correlation check (ρ = −0.4045) between detected PII count and estimated k-anonymity to validate the estimator’s monotonic behavior.

  • The improved system can flag when the estimator becomes uncalibrated (e.g., due to outdated census data) and trigger fallback to more conservative anonymity estimates or web-grounded retrieval for current population statistics.

  1. Multilingual and Scalable Privacy Pipeline (Future-Ready)
  • Extend the disclosure extraction and k-anonymity estimation to non-English languages by training on multilingual corpora and using language-specific population priors.

  • The improved system can protect users across global deployments, automatically adapting to local demographic statistics and cultural norms around self-disclosure.

What the Improved AI System Can Do:

  • Act as a privacy-preserving intermediary between users and powerful remote LLMs, rewriting queries to minimize identifiability while maintaining high response quality.

  • Detect and mitigate subtle self-disclosures that standard PII detectors miss, reducing risks from data breaches and training-data memorization.

  • Provide real-time, quantitative anonymity scores for user inputs, enabling transparent privacy decisions (e.g., “This query narrows you to 1,000 people”).

  • Optimize its own behavior across different local models and deployment constraints, ensuring consistent privacy-utility trade-offs.

  • Scale to large-scale conversational platforms, handling thousands of queries per second with cached, low-cost anonymity estimation.

Abstract

Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors. One such approach is Privacy-Conscious Delegation (PCD), where a local LLM acts as an intermediary. However, privacy risk does not stem solely from explicit identifiers but also PII-free self-disclosures, leaving users identifiable through combinations of quasi-identifying traits. We investigate a probabilistic variant of PCD, where we augment its objectives with an LLM-driven probabilistic estimation of k-anonymity. To facilitate this, we first create the PUPA-SD dataset, which contains naturalistic user queries with self-disclosure. Our preliminary results indicate that optimizing PAPILLON on PUPA-SD improves quality on unseen conversations across a variety of local models and produces the best privacy-utility balance for Llama-3.2-3B, while smaller models struggle to jointly optimize quality and privacy. We propose k-anonymity as a useful auxiliary metric for tackling PCD.

Sources

Related papers