SoK: From Generation to Consumption of Privacy Documents in Software Systems

arXiv:2608.12511 · cs.CR, cs.CL · Submitted 2026-08-12 · Read on arXiv

Shidong Pan, Clark LaChance, Zhen Tao, Sepideh Ghanavati

Columbia University · New York University · University of Maine · Technical University of Munich

cs.CR, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-14

Comments: This SoK paper has been accepted by NDSS 2027

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This SoK paper provides a unified, lifecycle-oriented view of privacy documents from a software engineering perspective.

Terminology

Summary

This SoK paper provides a unified, lifecycle-oriented view of privacy documents from a software engineering perspective. It systematically reviews and analyzes 290 papers published between 2010 and 2025, organizing them around five research questions that examine how privacy documents are (1) defined and scoped, (2) generated, (3) analyzed and extracted, (4) checked for inconsistencies and noncompliance, and (5) evaluated and improved for usability. Building on the findings, the paper identifies 15 key research trends and 21 open opportunities, and charts four broader research directions.

The paper's abstract states: "This SoK provides a unified, lifecycle-oriented view of privacy documents from a software engineering perspective. We systematically review and analyze 290 papers published between 2010 and 2025, organizing them around five research questions that examine how privacy documents are (1) defined and scoped, (2) generated, (3) analyzed and extracted, (4) checked for inconsistencies and noncompliance, and (5) evaluated and improved for usability. Building on our findings, we identify 15 key research trends and 21 open opportunities. We further chart four broader research directions that highlight (i) emerging challenges in AI-centric platforms, (ii) the need for diverse and up-to-date data foundations, (iii) LLM-based unified policy-code analysis, and (iv) dual usability for end-users and developers."

The introduction explains the motivation: "Privacy regulations commonly mandate that software systems provide user-facing disclosures that explain how personal data are collected, used, shared, and protected. These disclosures take the form of privacy documents, such as privacy policies, privacy labels, permission rationales, and cookie notices. Privacy documents are intended to serve as a central mechanism for transparency, accountability, and user empowerment across modern digital ecosystems. They also play a critical compliance role, acting as the interface between regulatory requirements, software behavior, and user understanding. Despite their importance, privacy documents often fail to achieve their intended goals. Prior work has repeatedly shown that privacy policies are lengthy, vague, and difficult to understand. More recent formats, such as privacy labels and contextual privacy policies, aim to address these shortcomings, yet they introduce new challenges related to accuracy, consistency, and maintainability. Empirical studies further reveal widespread mismatches between privacy documents and actual software behaviors, as well as inconsistencies across different document types, platforms, and jurisdictions."

The paper notes the fragmentation in the field: "Research on privacy documents has grown substantially over the past two decades and spans multiple disciplines, including computer science (CS), law, public policy, economics, and psychology. Within CS, research aims to address a broad range of problems, such as automated generation and natural language analysis of privacy policies, regulatory compliance checking, and usability-driven redesigns of privacy notices. However, this body of work remains fragmented, particularly across research areas such as privacy, software engineering, and natural language processing (NLP). This fragmentation is also reflected in the current survey literature. Existing surveys typically focus on narrow slices of the literature, such as NLP techniques for privacy policies or usability studies of specific notice designs. There is currently no systematization that unifies these efforts across document types, technical approaches, and application goals."

The methodology section describes the systematic review process: We followed formal SLR guidance by using predefined RQs, conducting systematic query-based searching and screening, and inductive open-coding. The data collection involved identifying 32 venues (including privacy/security conferences, software engineering conferences, NLP conferences, and journals), conducting keyword searches, filtering, and forward and backward snowballing. The final analysis included 290 papers.

The paper's key findings are organized by research question:

RQ1 (Format and Scope): The paper finds that privacy policies represent the earliest and still most widespread form of privacy documentation (T1), appearing in 227 papers. It also notes that The formats of privacy documents constantly evolve (T2), with privacy labels being the second most studied format (28 papers), and other forms like contextual privacy policies, privacy icons, runtime privacy notices, and data processing agreements appearing less often. The research is English- and mainstream-market centric (T3), with only a few exceptions from other linguistic and cultural contexts.

RQ2 (Generation): The paper finds that Generation is relatively less discussed (T4) and existing research mainly focus on artifact-driven Code-based Generators and Policy Summary Generation (T5). Code-based generators use program analysis and IDE integration to infer privacy disclosures from software artifacts. Policy summary generation digests long policies into brief summaries. The paper notes that Individual human-factors are commonly considered, whereas the collaborative and dynamic nature of software development environments is often overlooked (T6).

RQ3 (Analysis): The paper categorizes analysis methods into rule-based and learning-based approaches. It finds that Rule-based analysis methodologies provide structured representations and constitute the foundation of privacy document analysis (T7), including symbolic NLP, policy formalization, and manual coding-based analysis. It also finds that Learning-based analysis inherits the taxonomy from rule-based analysis (T8) and These techniques have evolved alongside advances in NLP; however, limited explainability and hallucinations continue to challenge practitioners (T9). The paper notes that Manual analysis methods remain a major approach in privacy document analysis, despite the rise of learning-based methods (T10).

RQ4 (Consistency and Compliance): The paper finds that Software-Policy inconsistency studies mainly focus on mobile applications, driven by the availability of mature program analysis frameworks (T11). It also finds that Existing studies mainly focus on mainstream regulations, namely GDPR (T12). The paper notes that Inconsistencies across sources are prevalent, such as mismatches between privacy policies and privacy labels (T13).

RQ5 (Usability): The paper finds that Privacy documents are widely suffer usability and readability issues, they are commonly evaluated using metrics such as length, Flesch Reading Ease Test scores, and measures of vagueness (T14). It also finds that Privacy disclosures are increasingly designed to be just-in-time, contextualized, and more readable (T15), with four recurring design paradigms: label-ization, machine-readable formats, contextualization, and personal privacy assistants.

Based on these findings, the paper identifies four broader research directions:

D1: AI-centric platforms and their emerging privacy challenges. The paper states: "While existing studies have extensively examined privacy documents within traditional software such as websites and mobile apps, the rapid emergence of AI-centric platforms, such as LLM-based systems, agents, and modular 'skills', is fundamentally reshaping how software is developed and deployed, as well as how personal data is collected and processed."

D2: Update-to-date and diverse data foundations. The paper notes: "Foundational datasets, such as those collected and maintained in Usable Privacy Policy Project, have played a critical role in advancing privacy document research. However, there remains a need for updated and more comprehensive data infrastructures that better reflect the evolving landscape of privacy practices."

D3: LLM-based unified policy-code analysis. The paper states: "LLMs have considerable potential for widespread adoption in privacy policy analysis, because of their strong capabilities in natural language understanding and their ability to capture nuanced semantic meanings beyond rigid predefined taxonomies."

D4: Dual usability: end-users and developers. The paper notes: An evident imbalance persists in usability research, which focuses on end-users while largely neglecting the usability challenges faced by developers.

The paper concludes: "Over the past decades, research on privacy policies and documents has expanded significantly. This SoK systematically organizes 290 papers published between 2010 and 2025 to provide a unified, lifecycle-oriented view of privacy documents from an engineering perspective. Specifically, we examine how privacy documents are defined and scoped, created, analyzed and extracted, assessed for inconsistency and noncompliance, and evaluated and improved for usability. Across these lifecycle stages, we summarize 15 research trends, identify 21 open opportunities, and chart four broader research directions. By synthesizing prior work from a lifecycle perspective, we hope this SoK serves as a shared foundation for future research on privacy documents and supports the development of privacy communication mechanisms."

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. AI-Powered Privacy Document Generator with Code-Artifact Integration
  • Improvement: Train an LLM to generate privacy policies, labels, and permission rationales directly from source code, API calls, and data-flow graphs (extending T5’s code-based generators).

  • Capability: Automatically produce accurate, context-specific privacy disclosures that reflect actual software behavior, reducing manual effort and human error. The system can update documents in real-time as code changes, addressing T6’s gap in collaborative/dynamic development.

  1. LLM-Based Unified Policy-Code Consistency Checker
  • Improvement: Develop a cross-modal model that jointly analyzes natural-language privacy documents and executable code (or static analysis output) to detect mismatches (T11, T13).

  • Capability: Automatically flag inconsistencies between a privacy policy and actual data collection/sharing practices in mobile apps or web services, and between different document types (e.g., policy vs. label). This reduces reliance on manual audits and supports compliance with GDPR and other regulations (T12).

  1. Explainable AI for Privacy Document Analysis
  • Improvement: Enhance learning-based NLP models (T8, T9) with attention visualization, rationale extraction, and confidence scores for each extracted privacy practice.

  • Capability: Provide auditors and developers with transparent, verifiable outputs—e.g., highlighting the exact sentence in a policy that supports a data-sharing claim, or flagging hallucinated extractions. This addresses the explainability and hallucination challenges noted in T9.

  1. Multilingual and Culturally Adaptive Privacy Document Generator
  • Improvement: Fine-tune LLMs on diverse, non-English privacy corpora and legal frameworks (addressing T3’s English-centric bias).

  • Capability: Automatically generate or translate privacy documents that are linguistically and legally appropriate for different jurisdictions (e.g., GDPR, CCPA, LGPD), including localized readability adjustments and culturally relevant examples.

  1. Just-in-Time Contextual Privacy Assistant for End-Users
  • Improvement: Build an AI system that uses real-time context (e.g., user’s current app, data being accessed, location) to generate short, personalized privacy notices (extending T15’s contextualization paradigm).

  • Capability: Deliver concise, readable, and actionable privacy information at the moment of data collection, improving user comprehension and empowerment (T14). The system can adapt complexity based on user literacy and preferences.

  1. Dual-Usability Developer-Focused Privacy Tool
  • Improvement: Create an AI-powered IDE plugin that assists developers in writing, maintaining, and aligning privacy documents with code (addressing D4’s neglect of developer usability).

  • Capability: Provide inline suggestions for privacy-relevant code comments, auto-generate policy sections from code changes, and alert developers to potential policy-code mismatches before deployment—reducing developer burden and improving compliance.

  1. Dynamic Privacy Document Updater for AI-Centric Platforms
  • Improvement: Design an LLM-based system that continuously monitors AI agents, modular skills, and data pipelines (D1) to update privacy documents as new capabilities or data flows emerge.

  • Capability: Automatically revise privacy policies and labels when an AI system gains new features, integrates third-party APIs, or changes data processing logic—ensuring ongoing accuracy and compliance in fast-evolving AI ecosystems.

  1. Unified Cross-Document Semantic Analyzer
  • Improvement: Train a model on multiple privacy document formats (policies, labels, cookie notices, DPAs) to create a shared semantic representation (T1, T2).

  • Capability: Automatically compare and reconcile inconsistencies across different document types and platforms (T13), and generate a single, unified view of an organization’s privacy practices for regulators and users.

  1. Readability and Vagueness Scoring Engine
  • Improvement: Develop an AI metric that goes beyond Flesch scores (T14) by using LLM-based semantic analysis to detect vagueness, ambiguity, and legal jargon.

  • Capability: Provide actionable feedback to document authors—e.g., “This sentence is vague; specify the exact data retention period”—and automatically rewrite sections to improve clarity while preserving legal accuracy.

  1. Privacy Document Evolution Tracker
  • Improvement: Use LLMs to analyze version histories of privacy documents and code repositories to identify drift and emerging risks (T2, T6).

  • Capability: Alert organizations when privacy practices change without corresponding document updates, or when new regulations require modifications—supporting proactive compliance and reducing noncompliance penalties.

Abstract

Privacy documents (e.g., privacy policies) are a central mechanism through which digital services disclose data practices and seek user consent. Over the past decades, research on privacy documents has expanded significantly, encompassing not only traditional privacy policies but also short notices (e.g., privacy labels) and interface-level transparency mechanisms. As this research area continues to grow, it has become increasingly difficult to obtain a coherent view of how privacy documents are created, analyzed, evaluated, and maintained across their lifecycle. This SoK provides a unified, lifecycle-oriented view of privacy documents from a software engineering perspective. We systematically review and analyze 290 papers published between 2010 and 2025, organizing them around five research questions that examine how privacy documents are (1) defined and scoped, (2) generated, (3) analyzed and extracted, (4) checked for inconsistencies and noncompliance, and (5) evaluated and improved for usability. Building on our findings, we identify 15 key research trends and 21 open opportunities. We further chart four broader research directions that highlight (i) emerging challenges in AI-centric platforms, (ii) the need for diverse and up-to-date data foundations, (iii) LLM-based unified policy-code analysis, and (iv) dual usability for end-users and developers. We hope this SoK provides a shared foundation for future research on privacy policies and privacy documents.

Sources

Related papers