Large-scale online deanonymization with LLMs
cs.CR, cs.AI, cs.LG
Submitted: 2026-02-18
Updated: 2026-02-25
Comments: 24 pages, 10 figures
Journal ref: 35th USENIX Security Symposium (USENIX Security 26), pp. 1947-1966 (2026)
License: http://creativecommons.org/licenses/by/4.0/
The gist: We show that large language models can be used to perform at-scale deanonymization.
Terminology
Abstract
We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each containing unstructured text written by or about that individual, we implement a scalable attack pipeline that uses LLMs to: (1) extract identity-relevant features, (2) search for candidate matches via semantic embeddings, and (3) reason over top candidates to verify matches and reduce false positives. Compared to classical deanonymization work (e.g., on the Netflix prize) that required structured data, our approach works directly on raw user content across arbitrary platforms. We construct three datasets with known ground-truth data to evaluate our attacks. The first links Hacker News to LinkedIn profiles, using cross-platform references that appear in the profiles. Our second dataset matches users across Reddit movie discussion communities; and the third splits a single user's Reddit history in time to create two pseudonymous profiles to be matched. In each setting, LLM-based methods substantially outperform classical baselines, achieving up to 68% recall at 90% precision compared to near 0% for the best non-LLM method. Our results show that the practical obscurity protecting pseudonymous users online no longer holds and that threat models for online privacy need to be reconsidered.
Sources
- LLMs unlock new paths to monetizing exploits
- Automated Profile Inference with Language Model Agents
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Gemini Embedding: Generalizable Embeddings from Gemini
- Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset
- Anonymity and Identity Online
- Position: Privacy Is Not Just Memorization!
- OpenAI GPT-5 System Card
- Adversaries Can Misuse Combinations of Safe Models
- RAT-Bench: A Comprehensive Benchmark for Text Anonymization
- On the State of the Art in Authorship Attribution and Authorship Verification
- A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs