MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-Ended Question Answering

arXiv:2603.14265 · cs.CL, cs.MA · Submitted 2026-03-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-Ended Question Answering".

Jane: The paper was written by author1 and author2 from University1 and Company2.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Jane: The authors of "MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off..." have set up this entire evaluation to address what they call contextual leakage, which is such a subtle and dangerous concept.

Tom: Contextual leakage—that's when the specific combination of details in a patient’s history allows them to be re-identified, even if you stripped out the obvious stuff like their name or address.

Lu: It’s not just about the explicit identifiers anymore; it’s about that unique constellation of medical facts that makes the identity possible, and this is something traditional de-identification methods fail to catch.

Meng: From an engineering standpoint, this means we can't just rely on basic scrubbing techniques because the semantic relationships within clinical narratives are far more complex than simple token removal.

Lalam: It’s a necessary shift in our thinking, recognizing that true privacy requires understanding the narrative structure of how a patient’s story is told.

Tom: And Jane, it's encouraging that we finally have a framework to quantify this threat, which the authors are doing by setting up this entire benchmark to quantify exactly what they mean by the privacy-utility trade-off.

Jane: It’s a very important distinction for establishing safety, and it makes me think about how far we have come from just thinking about generic hallucination problems.

Lu: We're moving toward understanding the real danger that contextual leakage poses to patient autonomy and compliance with regulations like HIPAA, which is a huge step forward.

Meng: It’s definitely something that needs to be rigorously tested, and I think this framework gives us the tools to do just that.

Lalam: We can't ignore these risks when we are building the future of healthcare, and understanding this trade-off is the first step toward building trustworthy AI systems.

Summary: Tom: So, we've established that contextual leakage is a problem, but what did MedPriv-Bench actually do to model this risk? I think the core of the system they built is fascinating.

Jane: They created a multi-agent pipeline to simulate this risk in a realistic way, which is far more sophisticated than just running random queries against an LLM.

Lu: The authors designed three distinct agents—the PHI Injection Agent, the Question Agent, and the Answer Agent—to fully replicate the entire workflow of how this leakage could happen in a real-world RAG system.

Meng: The injection agent takes de-identified chunks and rewrites them by embedding specific, sensitive facts from a taxonomy, making sure those fragments are relevant to the medical narrative.

Lalam: It’s quite clever that these agents are designed to make the injected PHI not just random secrets but details that would actually improve the quality of an answer if they were used.

Tom: And that's where it gets tricky, because the question agent then crafts a query based on those newly added sensitive facts, making sure the resulting query is clinically meaningful.

Jane: It’s a complete simulation of the threat model, which is great because it shows us exactly how a white-box attacker might use this information against the LLM.

Lu: The process is designed to show that this leakage isn't just an accident; it’s a logical consequence of using specific details in combination with medical queries.

Meng: I find the workflow really practical, because it gives us a defined path from raw data chunks to the final generated output, making it easy to track where the failure occurs.

Lalam: We are essentially observing how this sophisticated process moves us toward a future where we can understand and manage the complexity of AI in clinical decision support.

Improvements: Tom: Now that we understand how they build these tests, let's talk about the improvements in *how* they measure them, because that's another huge part of MedPriv-Bench. It’s not just about looking at the answer; it’s about measuring leakage.

Jane: They introduced a standardized evaluation protocol using a pre-trained RoBERTa-Natural Language Inference model as an automated judge, which is a massive leap from relying solely on human experts.

Lu: This NLI approach allows us to quantify data leakage by determining if the LLM’s generated response entails any of the specific injected PHI facts, giving us a quantifiable measure of risk.

Meng: The authors developed two distinct metrics: the Instance-Level Leakage Rate and the Fact-Level Leakage Rate, which provides a very granular way to assess both how often leakage occurs and how much sensitive information is exposed overall.

Lalam: This level of detail allows us to move beyond just seeing if an answer is right; we can now see *why* it might be unsafe, which elevates the conversation about AI quality in healthcare.

Tom: It's a much more rigorous way to approach the problem, and Jane’s right, it helps us quantify this trade-off that was previously hard to measure with tools like BLEU or ROUGE scores.

Lu: The alignment of this automated judge—achieving eighty-five point nine percent alignment with human experts—is a huge validation of the methodology itself, showing that our measurement system is reliable and trustworthy.

Meng: From an engineering perspective, this standardized measure gives us a scalable way to compare different LLMs against each other without having to rely on subjective human scoring for the entire dataset.

Lalam: We are establishing metrics that allow us to build a culture of safety, where the cost of privacy is clearly defined alongside the utility of AI.

Conclusion: Tom: So, we’ve seen how MedPriv-Bench models and measures this risk, but what does all this research tell us about the future? What’s the big picture implication?

Jane: It confirms that there is a pervasive privacy-utility tradeoff across all major LLMs evaluated, meaning that maximizing helpfulness often comes at the expense of strict data protection.

Lu: The paper is showing us that we cannot simply assume AI will be safe; it highlights the critical need for domain-specific benchmarks to validate safety in privacy-sensitive environments like healthcare.

Meng: My takeaway is that if we are deploying these models, we have a clear roadmap now to measure and mitigate these specific leakage risks using the tools provided by this framework.

Lalam: It encourages us to adopt a more cautious and responsible approach, ensuring that our pursuit of clinical efficacy does not lead to privacy compromise in a world where patient trust is everything.

Tom: I think we’re all excited about how much clearer the path is now, Jane. We have this new standard for accountability with MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-Ended Question Answering.

Lu: It provides a clear framework for future research to push us toward better balance.

Meng: I’m ready to start implementing these safety checks now, too.

Lalam: We are so glad to see this is helping us build the right systems for the culture of healthcare we want to see.

Tom: Thanks again for joining us all on this topic, and I hope you have a safe and wonderful week ahead!

University1 · Company2

cs.CL, cs.MA

Submitted: 2026-03-15

Updated: 2026-08-25

Importance score: 88/100

The gist: MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering Abstract and Motivation The paper addresses a critical, yet overlooked,

Key concepts

Contextual Leakage
This dangerous concept occurs when the specific combination of details in a patient's history allows them to be re-identified, even if obvious identifiers like names or addresses are removed. It involves unique constellations of medical facts.
Privacy-Utility Trade-off
This refers to the balance between maximizing a Large Language Model's helpfulness (utility) and ensuring strict data protection (privacy). The research aims to quantify how improving one often compromises the other.
MedPriv-Bench
This is a benchmark developed to evaluate LLMs by quantifying the privacy-utility trade-off. It uses a multi-agent pipeline and standardized metrics to simulate and measure potential data leakage risks in medical contexts.

Terminology

Summary

MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering

Abstract and Motivation

The paper addresses a critical, yet overlooked, privacy threat in healthcare AI systems: contextual leakage. Unlike explicit Protected Health Information (PHI) which is typically scrubbed by de-identification techniques (like NER), contextual leakage occurs when an LL leverages a patient’s rich clinical history to synthesize a response that contains a unique constellation of medical details that enables re-identification, even if the response is devoid of explicit identifiers. This subtle threat violates regulations such as HIPAA and GDPR. Current benchmarks in healthcare heavily focus on accuracy, ignoring this privacy issue. To fill this gap, the authors present MedPriv-Bench, the first benchmark specifically designed to jointly evaluate privacy preservation and clinical utility in medical open-ended question answering.

The MedPriv-Bench Framework

The framework utilizes a multi-agent, human-in-the-loop pipeline to synthesize sensitive medical contexts and clinically relevant queries that create realistic privacy pressure. The methodology involves four distinct phases:

  1. PHI Injection Agent (A inj): Since the source data is deidentified, this agent samples specific PHI entities (E phi based on the context C i and a predefined taxonomy T) and embeds these entities into the context chunks to create a sensitive context (C i*).

  2. Question Agent (A quest): Acting as a "white-box attacker, this agent generates a natural language query (q) that is both clinically meaningful and designed to be better answered by leveraging the injected PHI. The agent also generates a reasoning path (r q explaining the logical connection between the question and the sensitive attribute.

  3. Answer Agent (A ans): This agent synthesizes the gold-standard response (a gold) using both chunks and the injected PHI, ensuring that it protects privacy while providing useful medical guidance.

  4. Auto-Validation and Filtering: The system uses a rigorous filtering pipeline, utilizing GPT-5.2 as an automated Clinical Validation Agent to perform binary (Pass/Fail) quality audits based on two primary dimensions: Correctness (Non-Contradiction) and Completeness (Minimum Viable Utility).

** Evaluation Metrics**

The authors established a standardized, scalable protocol for privacy evaluation using a pre-trained RoBERTa-Natural Language Inference (NLI) model.

  • ** Privacy Leakage Detection:** The leakage detection task is framed as an NLI problem:
  1. Premise (P): The model response i.

  2. Hypothesis (H i,j): The injected PHI fact e i,j.

The NLI model classifies the relationship between P and H i,j into one of three categories: Entailment, Neutral, or Contradiction. A binary leakage indicator (I(i, e i,j)) is 1 if the NLI indicates Entailment.

  • ** Leakage Metrics:**

  • Instance-Level Leakage Rate (LR inst): Assesses the safety of the model responses by calculating 1/N where N is the total number of instances, and an instance is considered leaked if it reveals at least one injected PHI fact.

  • Fact-Level Leakage Rate (LR fact): Measures the total proportion sensitive information extracted by the attacker across the the entire dataset, calculated as Total Leaked Facts / Total Injected Facts.

For utility, a response receives a score from 0 to 5 based on how many key points (defined in a scoring rubric) are correctly covered.

** Results and Findings**

The evaluation of 9 representative LLMs (including open-source, closed-source, and medical-specific models) demonstrated a pervasive privacy-utility trade-off.

  • The Over-Reasoning Trap: An analysis revealed significant privacy risks stemming from advanced reasoning capabilities, particularly in closed-source models like GPT-5.2. This model achieved the highest utility score (4.91/5) but also demonstrated a high vulnerability, with a 58.4% LR inst when operating without defenses. The authors attribute this to an overreasoning effect: the capacity to synthesize complex details inadvertently becomes a liability.

  • Impact of Defense: The application of prompt-based protection (a lightweight intervention) successfully reduced leakage rates across all models, confirming its effectiveness. However, this privacy enhancement consistently came at a cost, as most models exhibited a decrease in response utility.

  • Leakage Source Analysis: A sensitivity analysis revealed that leakage is primarily a generation-stage issue. Furthermore, imposing a 100-word length constraint while using full context caused LR inst to plummet from 39.4% to 5.7% while preserving high clinical utility, suggesting that enforcing brevity is an effective proxy for HIPAA’s Minimum Necessary Standard.

  • Clinical Intent: The study also found that certain types of queries, such as "Infection Control & Wound Care, had the highest leakage rate (40.00%), compared to Medication Adherence & Management" (24.10%).

Conclusion

The findings underscore the necessity of domain-specific benchmarks to validate the safety and efficacy of medical AI systems in privacy-sensitive environments, providing a standardized methodology for validating medical LLMs’ safety and regulatory compliance.

Improvements for AI systems

Improvement: The current benchmark systems treat PHI leakage detection as a binary pass/fail mechanism. I propose integrating a Contextual Sensitivity Scoring (CSS) Module. This module moves beyond simple redaction and instead quantifies the clinical necessity of disclosing or referencing specific pieces of sensitive information when generating an answer.

What the Improved System Can Do:

The system will calculate a risk score for every piece of PHI: Risk Score = Sensitivity Index times Relevance to Query - Actionability Gain.

  • If the score is high (i.e., the information is highly sensitive but not critical to answering the user's immediate question), the system will proactively generalize or abstract the data point.

  • If the score is low (i.e., non-critical, overly personal details), it will automatically suppress disclosure, even if it was present in the source text, thereby adhering strictly to a minimum necessary standard before drafting any response.


This ensures the generated advice is clinically robust and actionable (as shown in Example 2's reference answer) without ever exposing the underlying, prohibited PHI.

Sources

Related papers