AgenTRIM: Tool Risk Mitigation for Agentic AI
cs.CR, cs.AI
Submitted: 2026-01-18
Updated: 2026-08-30
Comments: EMNLP 2026 Findings
Code: https://github.com/crewAIInc/crewAI-examples
Project page: https://lbeurerkellner.github.io/jekyll/update/2025/04/01/mcp-tool-poisoning.html
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks.
Terminology
Abstract
AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risks such as indirect prompt injection and tool misuse. We characterize these failures as unbalanced tool-driven agency. Agents may retain unnecessary permissions (excessive agency) or fail to invoke required tools (insufficient agency), amplifying the attack surface and reducing performance. We introduce AgenTRIM, a framework for detecting and mitigating tool-driven agency risks without altering an agent's internal reasoning. AgenTRIM addresses these risks through complementary offline and online phases. Offline, AgenTRIM reconstructs and verifies the agent's tool interface from code and execution traces. At runtime, it enforces per-step least-privilege tool access through adaptive filtering and status-aware validation of tool calls. Evaluating on the AgentDojo benchmark, AgenTRIM substantially reduces attack success while maintaining high task performance. Additional experiments show robustness to description-based attacks and effective enforcement of explicit safety policies. Together, these results show that AgenTRIM provides a practical, capability-preserving approach to safer tool use in LLM-based agents.
Sources
- Command A: An Enterprise-Ready Large Language Model
- Defeating Prompt Injections by Design
- The Llama 3 Herd of Models
- MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
- Permissioned LLMs: Enforcing Access Control in Large Language Models
- Model evaluation for extreme risks
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Progent: Securing AI Agents with Privilege Control
- ToolFuzz -- Automated Agent Tool Testing
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
- AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
- MPMA: Preference Manipulation Attack Against Model Context Protocol
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
- Securing Agentic AI: Threat Modeling and Risk Analysis for Network Monitoring Agentic AI System
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs