Privacy in Personalized AI Is a System Property, Not Just a Model Property
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Privacy in Personalized AI Is a System Property, Not Just a Model Property".
Nadia: Individual model- or componentlevel analyses may not capture all privacy risks arising in personalized AI systems, motivating a system-level perspective on privacy.
Elias: First, who's behind it and why it matters.
Paper summary: Nadia: To wrap up the discussion on "Privacy in Personalized AI Is a System Property, Not Just Just a Model Property," the paper really pushes us away from thinking about privacy as something that belongs only to the model itself.
Elias: That’s right; it forces us to consider the entire user–system interaction and all of its information flows across components and over time as the unit we analyze.
Priya: It seems like this framework provides a cohesive basis for evaluating complex privacy risks that arise from how personalized AI applications operate in practice.
Nadia: The four interconnected leakage channels—data access, inferential, behavioral, and compositional leakage—combined with the proposed audit requirements give us a systematic way to look at these issues.
Elias: It really changes the way we approach system-level audits because it highlights that privacy loss can emerge in ways that are not obvious when you test components in isolation.
Priya: I think the real impact is guiding researchers and auditors toward checking those interaction trajectories and internal information flows rather than just looking at final outputs.
Nadia: So, we're moving toward a method where we evaluate how the system behaves over time and across different contexts, which seems like a necessary step for this type of technology.
Conclusion: Nadia: So, to wrap up this discussion, we're talking about the paper "Privacy in Personalized AI Is a System Property, Not Just a Model Property" and what that means for us as listeners today.
Elias: Yeah, I think it really challenges how we think about security and privacy in these systems by putting the focus on the whole interaction rather than just the math inside one model.
Priya: And from my perspective as someone who looks at how data actually flows, this paper’s main contribution is making that flow visible through those four leakage channels.
Nadia: Exactly; it moves us away from thinking about a single privacy guarantee on a model and toward evaluating the entire system's behavior over time.
Elias: I agree, the title itself is pretty direct in signaling that we need to look at the architecture and interactions together, not just the isolated algorithms.
Priya: What’s striking is how it connects those technical leakage concepts—data access, inference, behavioral—to real-world scenarios we see in personalized AI every day.
Nadia: It seems like this paper provides a solid framework for auditors to start asking the right questions about where and when privacy risks actually manifest in these applications.
Elias: And I’m curious if the authors suggest any specific ways we can mathematically formalize those system-level requirements, like how to prove a system is truly protected across all those channels.
Priya: That leads us into how we can practically measure these risks; it suggests that utility and privacy need to be assessed together from the start, which is a big shift in measurement methodology.
Nadia: Exactly; it’s not just about protecting data in a static snapshot, but understanding the dynamic process of information exchange within the AI system itself.
Elias: So, this paper really sets up a new standard for how we should be evaluating these complex personalized AI setups moving forward.
Guillaume Salha-Galvan, Jiaying Xu
SJTU Paris Elite Institute of Technology · Kibo Ryoku Research
cs.CR, cs.IR, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Comments: NeurIPS 2026 Workshop on Privacy in the Era of Large Opaque Models
License: http://creativecommons.org/publicdomain/zero/1.0/
Importance score: 88/100
The gist: Individual model- or componentlevel analyses may not capture all privacy risks arising in personalized AI systems, motivating a system-level perspective on privacy.
Key concepts
- Data-access leakage
- This occurs when private information the system has may be accessed or sent beyond what is necessary for the current task. For example, an assistant might retrieve medical records or a shopping agent could transmit a home address to an external tool without authorization.
- Inferential leakage
- Systems can learn sensitive details by combining patterns, even if the user never explicitly shared that information. This is relevant when the system derives knowledge beyond what is expected for its task and keeps it for future use or exploits it against the user's interests.
- Behavioral leakage
- This refers to what others can learn from the personalized outputs and actions of a system. Personalized recommendations, rankings, or advertisements can reveal private characteristics about the user, such as political affiliations or health status.
- Compositional leakage
- Privacy loss happens when multiple components or different interactions are considered together. An observer might link seemingly unrelated pieces of information—like a calendar entry and an old conversation—to uncover sensitive details about the user's life.
Terminology
Summary
Individual model- or componentlevel analyses may not capture all privacy risks arising in personalized AI systems, motivating a system-level perspective on privacy.
The gist
Privacy in personalized AI should be treated as a property of the system as a whole, with the full user–system interaction and its information flows across components and over time as the unit of analysis.
Personalization Expands the Privacy Surface
Personalized AI systems expand the privacy surface because system behavior depends on information collected, inferred, and reused about users.
Unlike standalone models, these systems may accumulate information across interactions, maintain explicit or implicit user representations, retrieve information when needed, route information through external tools, and adapt its outputs accordingly.
Privacy risks arise not only from what is directly observed but also from what the system accesses, infers, exposes through its behavior, and combines over time.
Privacy-Risk Channels In Personalized AI
The paper distinguishes four interconnected privacy-risk channels to characterize how privacy loss manifests:
-
Data-access leakage: This occurs when
private information available to the system may be accessed or transmitted beyond what is necessary, expected, or authorized for the current task.
Examples include an assistant retrieving medical information or a shopping agent transmitting a home address to an external tool. -
Inferential leakage: A system can
infer sensitive information
by combining patterns, such as a recommender system inferringpolitical interests even if the user never declared them.
This becomes privacy-relevant when the system derives knowledgebeyond what is necessary or reasonably expected for the intended task, particularly when that knowledge is retained, reused across contexts, or exploited against the user’s interests or expectations.
-
Behavioral leakage: This concerns what others can learn from personalized outputs and actions.
Rankings, recommendations, explanations, advertisements, or other actions may all expose private characteristics of the user,
such as a feed recommending content associated with a political movement. -
Compositional leakage: Privacy loss emerges when
multiple components, contexts, or interactions are considered together.
An observer might correlate information across contexts—for example, combining a calendar entry and an earlier conversation to reveal that a user is undergoing medical treatment.
Requirements for System-Level Privacy Evaluation
To guide system-level privacy evaluation in future privacy audits of personalized AI, the paper proposes four requirements:
-
Evaluate interaction trajectories, not only individual responses: Evaluations should include
longitudinal user histories and allow adaptive probing based on prior outputs
becauseseveral individually innocuous responses may jointly reveal sensitive information that no single response exposes.
-
Audit internal information flows, not only final outputs: Audits must examine
which information is retrieved, which tools or agents receive sensitive data, and what is retained afterward,
focusing on whether access isnecessary, authorized, and appropriate for the current task and context.
-
Do not neglect inferential and behavioral leakage: Evaluations should assess
what sensitive attributes the system can infer from user behavior
as well aswhat an external observer can learn about the user from the system’s personalized outputs or actions.
-
Assess privacy jointly with personalization utility: Evaluation must consider utility alongside privacy loss, recognizing that
unrestricted access to user data may improve personalization while increasing risks across the four channels,
and connecting this to data minimization:the relevant question is not only whether information is protected, but whether the system needs to access or retain it in the first place.
Discussion
The framework complements model-level analyses by describing how privacy loss manifests rather than focusing on mechanisms that trigger it. The four channels are interconnected; for instance, a private medical event can trigger data-access leakage, which leads to inferential leakage, which alters behavioral leakage, and is ultimately exposed through compositional leakage. System-level audits must be precise: the model is differentially private,
the model did not memorize the user’s data,
and the system protects the user’s privacy
are distinct statements. The paper stresses that personalization creates interconnected risks across components and over time, which isolated component evaluations cannot capture. Formal guarantees like differential privacy remain valuable but do not characterize risks arising when sensitive information is introduced through profiles or external tools, which is the scope of system-level audits.
Conclusion
The paper argues that privacy in modern personalized AI applications cannot be reduced to a model property and proposes a cohesive basis for system-level privacy evaluation by analyzing four interconnected leakage channels and proposing four corresponding audit requirements. These contributions aim to provide a systematic approach for evaluating the complex privacy risks inherent in personalized AI.
The gist
Privacy in personalized AI should be treated as a property of the system as a whole, with the full user–system interaction and its information flows across components and over time as the unit of analysis.
How it works
Personalized AI systems expand the privacy surface because "system behavior depends on information collected, inferred, and reused about users.
Improvements for AI systems
Based on the research presented in Privacy in Personalized AI Is a System Property, Not Just a Model Property,
here are specific improvements for building more robust and privacy-aware personalized AI systems:
The proposed system-level privacy framework suggests moving beyond isolated model audits to evaluate the entire user-system interaction. The following improvements target the four identified risk channels (data-access, inferential, behavioral, compositional leakage) through the proposed four requirements for system evaluation.
-
[Requirement: Evaluate interaction trajectories]
-
[Requirement: Audit internal information flows]
-
[Requirement: Do not neglect inferential and behavioral leakage]
-
[Requirement: Assess privacy jointly with personalization utility]
Here is what the improved AI system can do, categorized by the improvement focus:
-
[Based on Requirement 1 (Interaction Trajectories)] The system will be capable of
Adaptive Privacy Probing.
-
This means when a user asks a series of seemingly innocuous questions over multiple turns (e.g., discussing travel plans, then health concerns, then political views), the system will monitor the accumulation of these interactions to detect patterns that constitute
compositional leakage.
If the sequence reveals sensitive information (like medical treatment) that no single response would expose, the system will pause and request explicit user consent before proceeding with any subsequent personalized action based on that inferred knowledge. -
[Based on Requirement 2 (Internal Information Flows)] The system will implement a
Contextual Data Access Gate.
-
This means the system will not simply output information; it must log and justify every piece of data retrieved from user profiles, external tools, or memory components before use. If the system retrieves a sensitive item (e.g., home address for a shopping request) but does not explicitly state in its internal flow logs that this retrieval is necessary for the current task's objective (and authorized by policy), it will halt and require explicit authorization from the user or an administrator before transmitting or utilizing that data, preventing unauthorized
data-access leakage.
-
[Based on Requirement 3 (Inferential and Behavioral Leakage)] The system will incorporate a
Sensitivity Guardrail for Inferences.
-
This means the system will be trained not just to provide accurate recommendations but also to be calibrated against the risk of inferring sensitive attributes (like political interests or health status). If the personalization engine attempts to derive an attribute that is deemed highly sensitive (e.g., inferred political affiliation), it must flag this inference as a potential
inferential leakage
and refuse to use that inferred knowledge for further personalization unless a higher level of explicit user consent is provided specifically for that type of inference. -
[Based on Requirement 4 (Privacy-Utility Trade-off)] The system will employ
Utility-Aware Data Minimization.
-
This means the system will be architected to dynamically adjust its level of data access and inferential scope based on the immediate task's utility requirement versus the associated privacy risk. For example, if a user asks for a simple restaurant recommendation (high utility, low risk), the system will only use explicitly provided preferences and avoid accessing long-term interaction history. If a more complex task requiring deeper memory retrieval is requested, it will automatically trigger a higher privacy check to ensure that the necessary information is accessed minimally and that any resulting behavioral leakage is constrained to the immediate context of that task.
Abstract
In personalized AI applications, such as conversational assistants and recommender systems, users interact not with models in isolation but with broader systems that access, infer, and reuse user information across components and over time. While such use of user information is integral to personalization, it also raises important privacy questions. In this paper, we argue that individual model- or component-level analyses may not capture all privacy risks arising in such systems, motivating a system-level perspective on privacy. We distinguish and analyze four interconnected privacy-risk channels in personalized AI, and subsequently propose four requirements for system-level privacy evaluation, covering interaction trajectories, internal information flows, indirect leakage, and the privacy-utility trade-off. We argue for their systematic incorporation into privacy audits of personalized AI.
Sources
- On the Opportunities and Risks of Foundation Models
- Imprompter: Tricking LLM Agents into Improper Tool Use
- ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
- A LINDDUN-based Privacy Threat Modeling Framework for GenAI
- The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
- Position: Privacy Is Not Just Memorization!
- PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
- Scalable Extraction of Training Data from (Production) Language Models
- The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
- PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs