Context Inference Attacks Without Jailbreaks
cs.CR, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden context before the system answers.
Terminology
Abstract
Agentic AI systems are increasingly deployed to process sensitive data at inference time, such as healthcare records or financial documents assembled into a hidden context before the system answers. Prior work has studied privacy risks primarily through jailbreaking attacks that induce models to directly disclose sensitive content, but has largely overlooked the agentic setting where the context is assembled by the agent's own tool calls. We show that the agents we evaluate remain vulnerable to hidden-context leakage despite the controls we test against them, namely an instruction not to disclose the context, logit suppression, and context dilution. For instance, a web-browsing agent answering benign user queries still carries exploitable signals about records silently loaded into its context. We introduce and formalize context-inference attacks through a security game and evaluate three settings under decreasing attacker knowledge and increasingly indirect delivery of the context: a known context, an unknown context, and a context the agent retrieves through its own tool calls. We distinguish a grey-box setting, in which the target model is used to score observations, from black-box settings in which the attacker scores with a surrogate it controls. We further characterize how leakage varies with query budget, context size, and target-model size. A single attack carries through all three settings without modification, reaching 100% ASR on small candidate sets and 63% at 1024 candidates against a known context, 78.9 AUROC when the template and surrounding records are unknown, 92.5 AUROC when a 14B surrogate scores a 32B target, and 81.8 AUROC when the records arrive as an agent's retrieval returns, against chance rates of 1/Z and 50 respectively.
Sources
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
- Qwen2.5-VL Technical Report
- Beyond Memorization: Violating Privacy Via Inference with Large Language Models
- Evaluating Large Language Models Trained on Code
- On the Privacy Risk of In-context Learning
- Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
- Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
- Last One Standing: A Comparative Analysis of Security and Privacy of Soft Prompt Tuning, LoRA, and In-Context Learning
- Qwen2.5 Technical Report
- Effective Prompt Extraction from Language Models
- Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs