Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
cs.CR, cs.AI, stat.AP
Submitted: 2026-06-26
Updated: 2026-06-26
Comments: 15 pages, 2 figures
Journal ref: Privacy in Statistical Databases (PSD 2026), Lecture Notes in Computer Science, vol. 16925, pp. 319-332, Springer, Cham (2027)
DOI: 10.1007/978-3-032-37883-5_21
License: http://creativecommons.org/licenses/by/4.0/
The gist: The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public.
Terminology
Abstract
The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely recognised by the public. While prior research has established that mobility traces are highly unique and that individuals can, in principle, be identified from a handful of spatio-temporal points, such attacks have historically required significant manual effort from skilled analysts, limiting their practical scale. In this feasibility study, we demonstrate in a real world setting that agentic AI fundamentally changes this threat model. We present an end-to-end pipeline in which large language model agents autonomously search the open web, cross-reference public records and social media, and resolve raw coordinate sequences to candidate identities - without human intervention. We evaluate the pipeline on a spatio-temporal dataset containing simulated location points anchored at and around true home and work addresses, focusing on a high-risk disclosure scenario. Our results demonstrate that, from spatio-temporal data and public sources alone, our agentic AI successfully re-identified 18 of the 25 re-identifiable individuals (72%) and 18 of 43 cases overall (41.9%). We discuss implications for Statistical Disclosure Control (SDC) practice and outline the near-future escalation that data custodians and regulators must anticipate. De facto anonymity - an implicit foundation of SDC practice - is shifting. Agentic AI strengthens the case that re-identification is reasonably likely by any means under the GDPR Recital-26 standard, at costs of minutes-and-dollars per target.
Sources
- The Llama 3 Herd of Models
- From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
- Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
- Beyond Memorization: Violating Privacy Via Inference with Large Language Models
- Investigating Vulnerabilities of GPS Trip Data to Trajectory-User Linking Attacks
- Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs