AI Snitches Get Glitches: Towards Evading Agentic Surveillance
cs.AI
Submitted: 2026-06-24
Updated: 2026-09-04
Comments: The code and demo are available at https://aisec.cs.umass.edu/demo/ai-snitches-get-stitches/
Code: https://github.com/umass-aisec/ai_snitches_get_stitches
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI agents are now routinely entrusted with access to users' data and communications, operating with growing autonomy and low human supervision.
Terminology
Abstract
AI agents are now routinely entrusted with access to users' data and communications, operating with growing autonomy and low human supervision. This increasing reliance on AI agents introduces a novel privacy risk that we call agentic surveillance, wherein third-party-provided agents leverage their access privilege to monitor for specific user behaviors, compile a targeted report, and covertly deliver it via tools. Users under surveillance may have neither the ability to control nor awareness of what the agents do on their behalf. To study the surveillance capabilities of different LLMs, we construct SURVEILBENCH, a benchmark dataset comprising over 300 diverse surveillance scenarios across domains. We find that several LLMs, such as Gemini 3.1 Pro, report users in at least 3--30% of cases, even when they are not explicitly instructed to do so. Despite safety guardrails and alignment to protect user privacy, almost all models can be readily prompt-tuned to conduct extensive surveillance in >75% of cases. Intriguingly, we also observe the agents reporting the surveillance attempt itself to government authorities. Finally, we repurpose prompt injection for the opposite goal---evading surveillance---and develop three techniques that let users hide from, deceive, or induce over-escalation in surveillance agents. We conclude that agentic surveillance is already easy to implement in practice, and we call for a comprehensive technical, ethical, and legislative framework to protect users.
Sources
- AI Agents May Always Fall for Prompt Injections
- Why Do Language Model Agents Whistleblow?
- Constitutional AI: Harmlessness from AI Feedback
- Examining Risks in the AI Companion Application Ecosystem
- LLMs unlock new paths to monetizing exploits
- Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Large-scale online deanonymization with LLMs
- Can Large Language Models Really Recognize Your Name?
- Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies
- Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection