Extracting Knowledge from Tools in LLM Agents
cs.CR
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents commonly use knowledge-based tools and access their underlying files, databases, and search indexes through tool invocation.
Terminology
Abstract
LLM agents commonly use knowledge-based tools and access their underlying files, databases, and search indexes through tool invocation. This integration improves agents' ability to provide domain-specific services but also introduces the risk of tool-mediated knowledge extraction: source content exposed to an agent for legitimate responses may be progressively recovered from its outputs, enabling reconstruction of the knowledge source behind a target tool. This paper systematically investigates this risk and identifies two challenges introduced by tool invocation: tool-selection uncertainty, where an agent may invoke a competing tool instead of the target tool, and tool-argument compression, where fine-grained query information may be lost when the agent generates tool arguments. To tackle these challenges, we propose ToolSiphon, a query-only extraction attack that introduces two complementary signals: a target-discriminative signal, implemented through Tool Contrastive Analysis, to steer queries toward the target tool; and a response-grounded factual signal, implemented through Evidence Chained Feedback, to mitigate argument compression and progressively expand extraction coverage. Across three types of knowledge-based tools and six domain-specific datasets, ToolSiphon recovers 74.3% of source records on average when coarse-grained information about non-target tools is available, with 83.2% textual recovery and 90.2% semantic similarity. Even without such information, it recovers 66.3% of source records. ToolSiphon also remains effective against representative defenses and on three real-world agent platforms.
Sources
- RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
- LLM Agents for Education: Advances and Applications
- Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking
- ToolTalk: Evaluating Tool-Usage in a Conversational Setting
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
- ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
- Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
- Connect the Dots: Knowledge Graph-Guided Crawler Attack on Retrieval-Augmented Generation Systems
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs