Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning".
Jane: The paper was written by author1 and author2 from University1 and Company2.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: So, we’re talking about Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning. The core of this is that the system doesn't need human input to know what is private; it figures it out itself using context.
Jane: It's a big improvement over traditional static lists, right? The paper clearly shows that existing methods struggle because they either miss the sensitive details or they require someone to be an expert in a domain-specific privacy standard.
Lu: The "domain-aware" part is the key, and it speaks to how much richer the embedding space is than we’ve seen before. Instead of treating all medical symptoms as one thing, they are grouping them based on their underlying semantic meaning within the context of Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Meng: I'm interested in how this translates to real use; if the system is reliably identifying these domain-specific concepts, we can design a client-side module that is truly plug and play. This means less burden on the deployment side of things.
Lalam: And for Lalam, it represents a more natural way for AI to interact with complex human information; it's respecting the nuances of privacy without demanding human expertise from our users.
Tom: It's clear they are moving beyond just general text sanitization, and Jane, you’ can tell me what this sophisticated system is actually doing?
Summary: Jane: Essentially, the paper summarizes a two-phase approach to make privacy work reliably. The first phase involves building a knowledge base of what constitutes privacy across different domains. Then the second phase, which happens when you are actually using the AI, takes a user's query and surgically remove only those specific sensitive parts.
Tom: It’s not just masking random words; it’s finding segments that carry the actual confidential information. I remember seeing that distinction between "privacy spans" and regular context in the summary of Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Lu: And for me, the process of how they localize those spans is what’s intriguing—it's a dynamic inference rather than a static lookup, which allows the the system to handle unexpected or novel phrasing. It's not just looking for keywords; it's looking for semantic intent.
Meng: The "Online Inference" part of Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning is where I see practical utility; it chunks the text, figures out which domain we are in, and then pinpoint that the precise spans that need to be rewritten.
Lalam: For Lalam, this means a vision of an AI that has a high degree of situational awareness; it knows when it’s talking about medical issues versus legal disputes and can adapt its privacy protections accordingly.
Tom: It seems like they are solving the problem of "how to know what is sensitive" by making the system smart enough to infer that context.
Improvements/Mechanism: Jane: The major improvement, as detailed in Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning, is how they use these domain prototypes. They aren't just using a general model; they are using "Domain Privacy Prototypes" to guide the entire process.
Tom: That sounds like a very focused way to define what's sensitive—a set of semantic anchors that capture diverse examples within Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Lu: And I think the mechanism for creating those prototypes is quite clever, using multi-domain contrastive learning to ensure the embeddings are highly discriminative against non-private text. It’s forcing separation in the embedding space to capture the nuances of Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Meng: The use of DPO, or Direct Preference Optimization, is also a huge win for practical application. Instead of needing human annotators to define a perfect rewrite, they are using a composite reward function to generate preference pairs that automatically balance privacy against utility in Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Lalam: And the fact that this is all happening locally on the client side with a DP mechanism is what I love about Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning; it means we can achieve a high degree of privacy without relying on trusting an external server.
Tom: It seems like they are achieving that perfect balance where the system isn't just hiding words, but intelligently rewriting them to maintain semantic integrity.
Conclusion: Jane: So, looking at Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning in its entirety, we see a method that solves the problem of "untrustworthy" LLMs by putting the intelligence right on the client side.
Tom: It's a comprehensive approach that gives us much more control than just relying on static lists or brittle prompt engineering. The results are quite strong, showing significant performance gains across all metrics.
Lu: I think this work proves that sophisticated semantic representations can effectively replace rigid, keyword-based detection systems in the realm of data privacy. It' shows a huge potential shift toward continuous learning for future research needs.
Meng: From an implementation standpoint, the computational cost is very manageable, which is a massive win for deployment; we don't need to worry about this system bogging down the user experience while running Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning.
Lalam: The impact of this work, Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning, is that it allows AI to be used in high-stakes sectors with a level of trust that was previously impossible to achieve.
Tom: It's a truly remarkable paper; it's making privacy both automated and effective.
Jane: We're going to wrap up our discussion of Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning for today, but we are so excited to see what the next paper brings next time.
author1, author2
University1 · Company2
cs.CR
Submitted: 2026-08-23
Updated: 2026-08-25
Importance score: 84/100
The gist: The paper, titled "Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning," addresses the critical challenge of deploying Large Language Models (LLMs) in
Key concepts
- Domain-Aware Approach
- The system uses domain knowledge to understand the underlying semantic meaning of text. Instead of treating all data equally, it groups symptoms or concepts based on their specific context (e.g., medical vs. legal), making privacy extraction richer and more accurate.
- Prototype Learning
- This method uses 'Domain Privacy Prototypes'—semantic anchors—to guide the process of identifying sensitive information. It captures diverse examples within a domain, allowing the system to define what is private using focused semantic definitions.
- Online Inference
- Unlike static lookups, this dynamic inference method processes text by chunking it and determining the specific domain in use. This allows the system to pinpoint precise sensitive spans and handle novel or unexpected phrasing effectively.
- Direct Preference Optimization (DPO)
- DPO is a mechanism used to generate preference pairs that automatically balance privacy against utility. Instead of requiring human annotators for perfect rewrites, it uses a composite reward function to guide the system's rewriting process.
Terminology
Summary
The paper, titled Mask-Free Privacy Extraction and Rewriting: A Domain-Aware Approach via Prototype Learning,
addresses the critical challenge of deploying Large Language Models (LLMs) in privacy-sensitive domains—such as healthcare and finance—where transmitting sensitive queries poses significant risks. Existing client-side privacy rewriting methods are found to be insufficient; full-text methods often distort context, while span-level approaches rely on impractical manual masks or brittle static dictionaries,
and attempts to automate localization via prompt-based LLMs prove unreliable, as they suffer from unstable instruction following that leads to privacy leakage and excessive context scrubbing.
To overcome these limitations, the authors propose DAMPER (Domain-Aware Mask-free Privacy Extraction and Rewriting). DAMPER is designed to operationalize latent privacy semantics into compact Domain Privacy Prototypes via contrastive learning, enabling precise, autonomous span localization.
The methodology is structured in two distinct phases: Offline Training and Online Inference.
Offline Training Phase:
This phase focuses on building the necessary semantic anchors and aligning the rewriting policy.
-
Privacy Prototype Generation: Due to
domain-specific privacy semantics
(e)g., disease severity in medicine vs. crime classification in law, (f)), a multi-domain contrastive learning objective is employed to reshape the embedding manifold. The goal is to achieveintra-domain compactness among privacy spans while enforcing separability from non-privacy contexts and out-of-domain privacy spans.
This process generates Domain Privacy Prototypes (P k), which serve as astructured abstraction of diverse in-domain privacy expressions
(g). -
Preference Construction and Alignment: To manage the tension between obfuscation and context preservation, the authors use a preference learning framework guided by these prototypes. A composite reward function is used to evaluate candidate rewrites (Y(x)) against two criteria: Semantic Obfuscation (r priv), which measures how distant the rewritten spans are from the original sensitive information, and Domain Fidelity (r util), which measures the proximity of each rewritten span to its nearest domain prototype. The final score is r(y) = (1 - alpha) r priv(y) + alpha r util(y). This allows for the automatic construction of a preference dataset (D pref) without human annotations. The rewriter (pi theta) is then fine-tuned using Prototype-Guided Direct Preference Optimization (DPO), optimizing the policy to align these synthesized preferences while constraining deviation from a reference model (pi ref).
Online Inference Phase:
This phase functions as a plug-and-play module for domain-adaptive span localization and differentially private rewriting.
-
Text Segmentation: The user input is first decomposed into semantically coherent spans using the TextChunker.
-
Privacy Span Localization (Domain Inference): Since the target domain is unknown, the system uses the learned privacy prototypes to
jointly infer the global domain and localize sensitive spans.
For each span, a score d k,i is calculated based on its maximum cosine similarity to a prototype in each domain k. The global domain is inferred as the one with the highest average affinity. -
Localization Thresholding: The system then determines a dynamic threshold (gamma) by analyzing the distribution of affinities, ensuring that only spans exceeding this threshold are considered sensitive (x).
-
DP Span-level Rewriting: The rewriter applies the transformation only to the detected privacy spans, copying all non-private tokens verbatim. To provide rigorous protection, a sampling-based Exponential Mechanism is integrated into token generation. By bounding the utility sensitivity via coordinate-wise clipping and sampling with a temperature (tau 2), this mechanism achieves a per-token privacy cost epsilon token, leading to a total query budget epsilon text.
Key Contributions and Results:
The paper highlights that DAMPER significantly outperforms existing baselines, achieving a superior privacy-utility trade-off.
The framework demonstrates exceptional resilience: Table 3 shows that while general methods like DP-MLM suffer catastrophic degradation
when relying on ground-truth masks, DAMPER maintains stability. Furthermore, the system exhibits robustness to vocabulary scarcity (Figure 5), maintaining high recall even when restricted to only the Top-20% most frequent spans during training.
In summary, DAMPER provides a mask-free, client-side privacy rewriting pipeline that obviates manual annotation while offering a principled, span-restricted protection mechanism,
successfully decoupling sanitization from downstream LLMs.
Improvements for AI systems
Assume the role of a highly skilled domain expert whose work is critically important to the field.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
- DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
- Qwen2 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs