ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
cs.CR, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-18
Comments: 8 pages, 3 figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report.
Terminology
Abstract
Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure elsewhere in a multi-step session and lack common ground truth across outlets, reports, and tool paths. We introduce privacy exposure displacement, the mismatch between a local evaluation proxy and target-grounded session exposure, and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis. Across multiple enterprise-style environments and independently implemented runtimes, we observe three recurring patterns. An expected-outlet-only view misses 46.9% of exposure recovered by the visible-exit union; attacker self-reports combine omissions with high false discovery; and schema-aligned internal evidence usually precedes visible exposure at the request/probe level. Reducing model-visible returns changes this path but can eliminate normal-task success. Independent human review supports the adjudication pipeline while identifying harder console and candidate cases. These findings motivate benchmarks that declare the complete visible boundary, ground claims in pre-specified targets and authorization, and report privacy together with task utility.
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Jailbreaking Black Box Large Language Models in Twenty Queries
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Red Teaming Language Models with Language Models
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs