CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

arXiv:2608.09164 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li

University of Washington · University of Illinois Urbana-Champaign · Northeastern University

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted to COLM 2026

Code: https://github.com/PEACH-Research-Lab/CIDERhttps:

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment Published as a conference paper at COLM 2026 Authors: Bingcan Guo1, Eryue Xu2, Jijie Zhou3, Zhiping Zhang3,

Terminology

Summary

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Published as a conference paper at COLM 2026

1 University of Washington, 2 UIUC, 3 Northeastern University

Abstract

Aligning large language models (LLMs) with human privacy preferences requires capturing individuals’ disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user’s disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user’s disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.

1. Introduction

The paper addresses the challenge of aligning LLMs with individual privacy preferences. While existing work uses the Contextual Integrity (CI) framework to construct datasets capturing societal privacy norms, these norms are inherently coarse-grained and cannot account for individualized privacy behavior. The paper introduces CIDER (Contextual Information Disclosure Boundaries Elicited from Real Users), the first personalized dataset capturing individualized contextual disclosure behavior from 169 real users. CIDER consists of 1,650 contextual disclosure boundaries (14,850 human annotations) across 60 scenarios, instantiated into 320 distinct data-sharing contexts by varying communication roles and AI-mediated conditions. Each boundary represents an individual’s acceptable disclosure behavior in a given context, encoded as binary ratings across 9 disclosure variants that systematically vary along granularity and identifiability.

2. Task

The paper formulates a prediction task where models predict a user’s disclosure decision in a new interpersonal communication scenario given their historical disclosure behaviors.

2.1 Definitions

  • Contextual Disclosure Boundary: Privacy preferences refer to individuals’ situated judgments about whether and how private information should be disclosed. Grounded in Communication Privacy Management theory, privacy boundaries represent the behavioral manifestation of personal privacy preferences, shaped by situational factors such as perceived risk, privacy-utility trade-offs, social relevance, and AI involvement. The paper operationalizes privacy boundaries as contextual disclosure boundaries: structured representations of an individual’s disclosure decisions within a specific communication context.

  • Granularity & Identifiability: Granularity refers to the level of detail in a disclosure (General, Moderately detailed, Very detailed). Identifiability refers to the degree to which a disclosure includes personal identifiers (Not Identifiable, Partially Identifiable, Fully Identifiable).

2.2 Problem Formulation

Each user has a latent personal privacy preference p that is not directly observable. When a user encounters a communication scenario s, this preference manifests as a contextual disclosure boundary y, represented as a binary vector over a set of scenario-specific disclosure variants V(s). The task is to predict the contextual disclosure boundary ŷ for a new context cnew, given historical contextual disclosure boundaries B = (ci, yi) i=1 N.

2.3 Metrics

The primary evaluation metric is per-prediction accuracy, measuring the proportion of correct disclosure predictions across all (g, i) variants. Accuracy is averaged for each user across predictions.

3. CIDER Dataset

To construct CIDER, the authors designed and conducted an online study to collect real user data.

3.1 Material Preparation

  • Scenarios: 60 scenarios were selected from the PrivacyLens dataset, each specified by five contextual integrity attributes: data type, sender, subject, recipient, and transmission principle. Participants were assigned a communication role (sender, subject, or recipient).

  • Disclosure Variant Generation: Granularity and identifiability were operationalized using three levels for each dimension. For each scenario, nine disclosure variants representing all 3 × 3 combinations were generated using GPT-o3 with a four-step prompt.

  • Study Design: The study aimed to elicit participants’ contextual disclosure boundaries by asking them to indicate whether they felt comfortable with the information being shared in a particular way. An AI-mediated condition was introduced, where the study interface included a sentence about the data sender using their AI assistant to share the information.

3.2 Data Collection

The study was hosted on Qualtrics and deployed on Prolific in August 2025. Quality control measures excluded two responses for exceptionally fast completion times and removed 38 boundary sets due to a material error. This yielded 169 valid participant responses.

3.3 Dataset Summary

The dataset contains 1,650 contextual disclosure boundaries from 169 users for 60 interpersonal communication scenarios. Average yes rate decreases monotonically with increasing granularity and identifiability, ranging from 73.76% for G1-I1 to 29.21% for G3-I3. The “Yes” rate is 49.89%, yielding an approximately balanced label distribution.

4. Experiments

4.1 Experimental Set-Up

  • Models: 12 state-of-the-art open and proprietary LLMs were evaluated: Claude Sonnet 4.6, Llama 4 Scout, Llama 4 Maverick, Llama 3.1 8B, DeepSeek-V3.2, Qwen3.5-9B, Qwen3-32B, Qwen3-14B, Qwen3-8B, Ministral 3 8B, GPT-5.4 (with medium reasoning effort), and GPT-5.4 nano.

  • Prompts: Three personalization conditions were evaluated alongside a zero-history baseline:

  • HC: historical boundaries provided with full semantic context.

  • HL: historical boundaries provided without semantic context, but with two-dimensional variant labels.

  • H: historical boundaries provided without semantic context or variant labels.

  • No-history baseline: no historical boundaries provided.

4.2 Results

4.2.1 Boundary-Level Performance

  • Larger and reasoning models generally achieve stronger personalization performance. Under HC (k=6), GPT-5.4 reaches 72.45% and Claude Sonnet 4.6 reaches 71.98%.

  • More capable models better leverage semantic context for personalized disclosure prediction. Moving from HL to HC, GPT-5.4 improves by 3.76 pp at k=6.

  • More capable models benefit more consistently from additional personalization history. GPT-5.4 improves by 7.08 pp from k=1 to k=6 under HC.

4.2.2 Variant-Level Performance

Performance gains from personalization history are imbalanced across disclosure variants. Only Claude Sonnet 4.6 demonstrates consistent improvement, reducing both FP and FN rates across all variants. Other models exhibit heterogeneous and imbalanced error shifts, where reductions in one error dimension are often accompanied by increases in the other.

4.3 Error Analysis

The paper analyzes cases where per-user prediction accuracy is 0 at k=6. Among 433 completely incorrect predictions, 261 consisted entirely of false negatives, 64 consisted entirely of false positives, and 108 contained a mixture. Four representative error types were identified:

  • Self-Disclosure Bias: Models treat self-disclosure scenarios as socially normative, predicting all variants as acceptable.

  • Pattern Extrapolation: Models extrapolate dominant rating patterns without reasoning about semantics.

  • Norm-Based Override: Models prioritize contextual norms over user-specific preferences.

  • Rule-Based Generalization: Models derive simplified attribute-level decision rules and apply them without calibration.

4.4 Discussion

The findings highlight the value and challenge of modeling privacy preferences across diverse individuals and contexts. GPT-5.4’s individual-level boundaries tie or outperform group-based boundaries in 71.70% of cases and norm-based boundaries in 83.02% of cases. The paper also reveals a tension in the design space of privacy-preserving LLMs: frontier models better capture individual preferences, but their deployment often requires transmitting user data to centralized infrastructure, which itself is a significant privacy risk.

5. Related Work

The paper discusses prior work on privacy alignment of LLMs, including ConfAIde, PrivacyLens, and other benchmarks. It also covers Contextual Integrity benchmarks such as CI-Bench, PrivaCI-Bench, and CI-Work.

6. Conclusion

The paper introduces CIDER, a dataset capturing real users’ privacy preferences, comprising 1,650 contextual disclosure boundaries collected from 169 users across 60 interpersonal communication scenarios. Results show that inference-time personalization using in-context behavioral history improves personalized disclosure prediction. However, the effectiveness varies across models: larger reasoning models better leverage semantic context and behavioral history, whereas smaller models rely more heavily on structural cues. Only Claude Sonnet 4.6 consistently achieves reductions in both false-positive and false-negative rates across all variants.

Ethics Statement

The study was approved by the Institutional Review Board (IRB). All scenarios and sensitive information used in CIDER are hypothetical and do not correspond to real individuals or events. No personally identifiable information was collected or retained.

Reproducibility Statement

The CIDER dataset and visual card artifacts are released on Hugging Face (https://huggingface.co/datasets/peach-lab/CIDER), and code is available on GitHub (https://github.com/PEACH-Research-Lab/CIDER). Complete prompts and survey instruments are provided in the appendices.

Improvements for AI systems

Improvements to AI systems:

  1. Context-Aware Privacy Preference Inference: AI systems can be enhanced with a personalized privacy reasoning module that uses a user's historical disclosure decisions (e.g., 6 past examples) to predict their comfort level with sharing specific information (e.g., health data, location) across granularity levels (general vs. detailed) and identifiability levels (anonymous vs. fully identifiable). This enables the system to tailor its sharing suggestions to individual users rather than applying generic privacy norms.

  2. Balanced Error Correction for Privacy Decisions: AI systems can be trained to simultaneously minimize both false positives (over-sharing sensitive data) and false negatives (under-sharing benign information). Specifically, the system can adopt a calibration mechanism—inspired by Claude Sonnet 4.6's balanced performance—that adjusts its prediction thresholds per user and per disclosure variant, ensuring that improvements in one error type do not come at the cost of the other.

  3. Semantic Context Utilization for Small Models: Smaller AI models (e.g., 8B-parameter LLMs) can be augmented with a lightweight semantic encoder that converts raw scenario descriptions (e.g., sender shares medical test results with a spouse via AI assistant) into structured features (data type, relationship, AI mediation). This allows them to leverage semantic context as effectively as larger models, reducing their reliance on brittle heuristics like more detailed = less acceptable.

  4. User-Specific Boundary Memory: AI systems can maintain a persistent, per-user privacy profile that stores their contextual disclosure boundaries across scenarios. When encountering a new sharing request, the system retrieves the most similar historical boundaries (using similarity metrics on scenario attributes) and uses them as few-shot examples, improving prediction accuracy by up to 11.41 percentage points without requiring fine-tuning.

  5. Error-Type-Aware Fallback Mechanisms: AI systems can detect when they are likely to exhibit systematic biases—such as self-disclosure bias (assuming all self-sharing is acceptable) or norm-based override (ignoring user preferences in favor of societal norms)—by analyzing the pattern of their predictions (e.g., all yes responses). Upon detection, the system can switch to a conservative mode that asks for explicit user confirmation before sharing, or it can query the user for clarification on ambiguous cases.

  6. Privacy-Preserving Personalization Architecture: AI systems can be designed with a two-tier architecture: a local, on-device model that handles personalized privacy predictions using the user's historical data (never leaving the device), and a cloud-based general model that only receives anonymized, aggregated privacy patterns (not individual boundaries). This balances the accuracy gains of frontier models with the privacy risks of centralized data transmission.

What the improved AI system can do:

  • Proactively manage information sharing in interpersonal communications (e.g., messaging apps, AI assistants) by predicting whether a user would be comfortable sharing a specific piece of information (e.g., I have diabetes vs. I have a chronic condition) with a specific recipient (e.g., boss vs. friend), at a specific level of detail and identifiability, and adjusting its suggestions accordingly.

  • Adapt to individual privacy preferences over time by learning from a user's past decisions, even with limited data (as few as 6 examples), and applying this learning to new, unseen scenarios—such as a new workplace communication or a novel AI-mediated sharing context.

  • Avoid both over-sharing and under-sharing by maintaining a balanced error profile, ensuring that it does not leak sensitive details (e.g., full medical records) while also not withholding innocuous information (e.g., general health status) that the user would want to share.

  • Explain its privacy decisions by referencing the user's historical boundaries and the specific contextual factors (e.g., recipient relationship, AI involvement) that influenced its prediction, enabling users to understand and correct the system's behavior.

  • Operate in privacy-sensitive environments by keeping personal data on-device, while still benefiting from advanced reasoning capabilities via a federated or hybrid approach, thus aligning with user expectations of data confidentiality.

Abstract

Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.

Sources

Related papers