User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios".
Jane: The paper was written by Authors not found in provided text. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 2: Tom: Now that we know the setup, let's look at the summary of findings from "User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios." The core discovery here is a striking difference between human agreement and AI consistency.
Jane: Overall, participants were generally satisfied; ninety-two percent found the response completed the task, and most felt it respected privacy as well. This suggests that LLMs can be helpful while preserving privacy under specific conditions, which is a relief for many people.
Lu: But that initial positive trend doesn's tell the whole story, Tom. The users had a very low level of agreement when evaluating individual scenarios—a Krippendorff’s alpha of zero point three six—meaning people were quite divided on what was appropriate for those specific contexts.
Meng: That lack of consensus among humans is huge because it demonstrates how subjective these privacy decisions are; we're seeing that the interpretation of what's truly "private" changes based on individual context and internal norms.
Lalam: It suggests that for tasks like summarizing a meeting, there isn't just one correct answer, Lalam thinks. The human evaluation is accurately reflecting that inherent complexity of the real world where different perspectives clash.
Tom: And Jane, you noted how this lack of agreement makes the proxy LLMs look incredibly stable in comparison; it was a major finding showing how polarized the human responses were.
Jane: It’s almost as if the AI is too rigid, Tom, always converging on one single conclusion rather than reflecting that wide range of ways humans debate these matters and see different angles. This divergence is what forces us to rethink how we measure quality.
Paper discussion segment 3: Tom: Moving into the misalignment, which is perhaps the most critical takeaway from "User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios," this is where the study really shows its teeth.
Jane: It’s not just that users disagreed with each other; they also significantly disagreed with the AI judges, which is a serious concern for anyone relying on automation to assess these complex tasks.
Lu: The misalignment stems from how proxy LLMs fail to capture those subtle contextual nuances, missing the fine details that humans picked up on immediately based on their own life experience and familiarity with data.
Meng: If an engineer relies solely on those automated AI scores, they might miss the fact that multiple participants spotted a critical error or a sensitive leak in the response that the AI simply didn't even see at all.
Lalam: This is a cultural shift for us, Lalam thinks; we can't just assume consistency is good when human experience of privacy is inherently varied and nuanced across ninety-four different people.
Tom: And Jane, you mentioned how poor those correlations were between user average ratings and what the proxy LLMs gave?
Jane: The correlation was only weak to moderate, Tom; meaning AI scores do not reliably estimate what people actually think about privacy or utility in a real-world setting. This gap highlights the limits of relying on automated judgment alone.
Paper discussion segment 4: Tom: After seeing that misalignment, the authors are suggesting specific ways we should improve how we evaluate LLM-generated content, which is a crucial step forward for future development in this field.
Jane: They are strongly advocating for human-centered evaluation, which is a massive shift away from the "LLM as a judge" approach that has been standard in research for years. We need to prioritize real human judgment moving forward with this paper.
Lu: The paper suggests moving beyond just seeking one single score and trying to capture the entire spectrum of user evaluations, Lu thinks, which is incredibly powerful for creating diverse and robust AI systems tailored to specific needs.
Meng: From an engineering standpoint, this means we must develop a more complex taxonomy—a clear framework that separates objective tasks from these preference-sensitive scenarios to guide our design decisions in practice.
Lalam: It's about designing AI that is not only helpful but also understands the diverse ways in which people define privacy and utility, Lalam thinks, so the technology serves the human needs of society.
Tom: And Jane, you see how this approach can help guide research and metric selection when we are trying to be more rigorous about what counts as a successful AI response?
Jane: Absolutely; we need to move away from simple proxy LLM evaluations and toward recognizing that human diversity is essential for ensuring the AI’s output makes sense in the real world.
CONCLUSION: Tom: We’ve covered a lot of ground today, looking at the findings of "User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios." It’s clear that for tasks involving personal privacy, relying on automated AI judges isn't enough to get a full picture.
Jane: The human experience must be the final authority, and it seems the consensus is that this study gives us a clear map of where we need to focus our efforts in LLM evaluation.
Lu: I think this work opens up so many creative possibilities for building more sophisticated, context-aware AI systems that can accurately reflect human judgment in our future.
Meng: And I’m glad we are seeing a push toward a more robust evaluation framework that accounts for real-world user input instead of just relying on consistent AI output scores.
Lalam: It's about making sure the technology serves the human experience rather than dictating it, and that's what this paper helps us understand regarding our societal impact.
Tom: It really shows how vital it is to listen to human judgment in the age of advanced AI models, even if those models are highly consistent themselves.
cs.CL, cs.AI, cs.HC
Submitted: 2025-10-23
Updated: 2026-09-03
Code: https://github.com/pln-fing-udelar/fast-krippendorff
Importance score: 80/100
The gist: The paper investigates user perceptions of Large Language Model (LLM) responses when handling privacy-sensitive scenarios, contrasting these human judgments against evaluations made by several proxy
Key concepts
- Proxy LLM Judges
- These are automated AI systems used to score the quality of LLM responses. The study found these judges were highly consistent but lacked nuance, failing to capture subtle contextual details that humans immediately recognized in the real world.
- Human Agreement
- This measures how much human participants agreed on what was appropriate for specific scenarios. The study noted a low level of agreement (Krippendorff’s alpha of 0.36), demonstrating that interpretations of privacy are highly subjective and vary widely among individuals.
- Misalignment
- This is the critical gap where human evaluations significantly disagreed with the scores provided by the proxy LLM judges. The weak correlation between these automated scores means AI cannot reliably estimate what people actually think about utility or privacy in a real-world setting.
- Human-Centered Evaluation
- This is a proposed shift away from using 'LLM as a judge' methods. It prioritizes real human judgment and capturing the entire spectrum of user evaluations to ensure AI output reflects diverse human needs and societal context.
Terminology
Summary
The paper investigates user perceptions of Large Language Model (LLM) responses when handling privacy-sensitive scenarios, contrasting these human judgments against evaluations made by several proxy LLMs. This research is critical because it assesses the reliability and consistency of automated AI judging mechanisms compared to real-world user experience, providing insight into both the utility and ethical compliance of generative AI systems.
User Perceptions of LLM Responses
Participant evaluations reveal high levels of perceived success regarding task completion and adherence to privacy norms. Specifically, participants found that 78% of the time, participants found the LLM-generated response to comply with the privacy norms,
and Over 82% of the time, participants found the response respected their personal privacy preferences.
Furthermore, concerning overall utility, data indicates that participant ratings showed high confidence in AI output:
-
Participants found that LLM-generated responses completed the given task over 90% of the time.
-
The perceived helpfulness and likelihood of use were also extremely high, with participants finding the responses
helpful and would use the response most of the time.
LLM Consistency vs. Human Evaluation
When comparing human judgment against proxy LLMs (such as GPT-5, Llama-3.3, Mistral, and Qwen-3), significant differences in evaluation consistency emerge. Figure 7 illustrates that Proxy LLMs show a lower standard deviation than participants’ evaluations on more scenarios across the board,
suggesting greater internal reliability in their judgments compared to the variability observed among human participants. Regarding correlation with user perception, Table 3 notes that Proxy LLMs’ evaluations moderately correlate with participants’ average perceptions of privacy preservation.
Conversely, the correlation between participant perception and proxy LLMs was found to be weak for helpfulness.
Semantic Consistency and Demographic Profile
The comparison extends beyond simple agreement rates to measure the quality of the generated explanations. Table 4 reports on semantic similarity, finding that Proxy LLMs provided responses that were much more semantically consistent than participants.
This suggests that while proxy models may not perfectly mirror human judgment across all metrics, their output structure is highly reliable. The study also collected detailed demographic data from 94 survey participants (Table 5), providing a profile of the user base:
-
Age: The largest group was aged 35 - 44 (35.11%).
-
Gender: Participants were slightly skewed toward males (51.06% Female vs. 2.13% Prefer to self-describe).
-
Education: The most common educational attainment was Bachelor’s degree (30.85%), followed by Master’s degree (15.96%).
-
Income: The largest income bracket represented was 100,000 or more (37.23%).
Improvements for AI systems
Disclaimer: As a diligent AI researcher, I have analyzed the provided empirical data concerning LLM performance relative to human perception regarding privacy, helpfulness, and task completion. The findings highlight superior consistency and semantic reliability in proxy LLMs compared to human participants. My proposed improvements focus on transforming this observed consistency into predictable, auditable ethical reasoning within the AI architecture.
-
Improvement: Develop a dedicated, modular layer that processes any input query through a structured, multi-stage ethical filter before generating the final response text. This module must move beyond surface-level keyword compliance (e.g., simply stating
I respect your privacy
) to perform deep semantic analysis of data flows and potential harms. -
Technical Implementation: The ERM should be trained not just on compliant examples, but crucially, on counterfactual ethical scenarios (i.e., inputs that should violate privacy or fail the task) paired with detailed human-annotated justifications for the correct remediation.
-
What the Improved System Can Do:
-
Guarantee Predictable Ethical Output: The system can ensure that its response adheres to a mathematically defined ethical boundary, significantly reducing variance in privacy compliance (P3/P4) regardless of prompt phrasing.
-
Identify and Flag Ethical Vulnerabilities: It can proactively identify potential points of data leakage or privacy violation within the user's own prompt and advise the user on how to rephrase the query for greater safety, rather than waiting for a failure point.
-
Improvement: Instead of treating privacy compliance as a single binary score, the system must implement a dynamic mapping function that quantifies the relationship between the Sensitivity Level of Input Data (P1) and the Required Mitigation Strategy. This moves from correlation (rho) to causation.
-
Technical Implementation: Fine-tune the model using a hierarchical ontology of data types (e.g.,
Pseudonymized Health Data
to High Sensitivity to Requires Differential Privacy Masking). The system must generate an internalRisk Score
alongside the response, detailing which parts of the input triggered which level of mitigation. -
What the Improved System Can Do:
-
Provide Justifiable Transparency: When challenged on its privacy adherence, the system can generate a detailed audit trail explaining why it masked certain information (e.g.,
The name was generalized because it is a direct identifier, triggering Mitigation Level 3
). This directly addresses the need for explainability that current correlation measures do not provide. -
Adjust Helpfulness Based on Risk: It can modulate its helpfulness (H2) based on the risk score; if the risk is too high, it must refuse to complete the task entirely, stating that doing so would compromise ethical integrity, thus improving safety over mere utility.
-
Improvement: Leverage the finding that LLMs are
much more semantically consistent than participants
(Table 4) by integrating a mandatory source attribution and confidence scoring mechanism for every factual claim made in the response. -
Technical Implementation: Every generated statement must be linked back to a specific, verifiable conceptual or data segment within its training corpus or provided context. The model must output not just the answer, but also a Confidence Interval (CI) for that answer's veracity and ethical standing.
-
What the Improved System Can Do:
-
Eliminate Hallucination in Ethical Contexts: It prevents the generation of confidently stated, yet factually or ethically flawed, information. If the CI drops below a set threshold (e.g., 0.95) for any ethical claim, the system must flag that specific sentence as speculative and require human review before outputting it.
-
Facilitate Debugging and Auditing: For regulatory or legal purposes, this layer allows an external auditor to trace every part of the AI's reasoning back to its originating evidence base, satisfying the highest standards of scientific diligence.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
- Human-Centered Privacy Research in the Age of Large Language Models
- Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social Relationships
- Analyzing Privacy Policies Using Contextual Integrity Annotations
- Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
- Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
- AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering