Beyond Direct Access: Resource Hijacking in LLM Agents
summary
The gist
The gist Large language model agents can be exploited by attackers to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or
In short
The research introduced agent resource hijacking, an attack where malicious agents trick other LLM agents into using or controlling high-value resources—like computing power or credentials—for the attacker's goals without stealing them directly. A systematic benchmark was created to study this threat, revealing that current defenses are insufficient against this indirect exploitation.
Key concepts
- Agent Resource Hijacking
- This attack occurs when an attacker manipulates an agent into performing actions like consuming compute time or accessing data for the attacker's benefit. It differs from credential theft because the resource itself is used legitimately by the agent, making detection very difficult.
- High-Value Agent Resources
- These are critical assets agents interact with, categorized into six types: Material (hardware like GPUs), Condition (API keys), Energy (usage quotas), Social and symbolic status, Information (private code), and Interaction channels. These resources provide the agent with significant power.
- ResourceHijackBench
- This is a testing system used to study the threat. It automatically generates 300 attack scenarios based on six resource categories, creating concrete targets and prompts. This allows researchers to systematically test how agents handle various types of resource misuse.
Terminology used across episodes
This episode discusses
- Beyond Direct Access: Resource Hijacking in LLM Agents · Paper Radio
- Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework
- Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing · Paper Radio
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
- When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
- ReAct: Synergizing Reasoning and Acting in Language Models
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
The paper
Beyond Direct Access: Resource Hijacking in LLM Agents · Read on arXiv
College of Cryptology and Cyber Science, Nankai University · Shanghai Jiao Tong University
Large language model agents are increasingly connected to high-value resources, including external APIs, GPUs and servers, and workflows such as deployment and approval. Existing agent security research mainly focuses on attacks against information and agent behavior, while high-value resources have received less attention as attack targets themselves. To our knowledge, we are the first to identify and systematically study agent resource hijacking, in which attackers induce agents to use high-value resources for their own goals without directly obtaining those resources or their credentials. We introduce ResourceHijackBench, an executable benchmark and automated case-generation pipeline that covers six categories of high-value resources. It contains 300 attack scenarios and 900 attack prompts, and each case runs in an isolated local environment that records actual resource use. Resource hijacking remains effective across four model backends, with average attack success rates of 70.0% to 89.6%, and it also appears across two agent harnesses, with average ASRs of 84.1% on OpenClaw and 72.3% on Codex. In paired comparisons, resource hijacking achieves ASRs 62.4 to 84.0 percentage points higher than direct resource acquisition across the four model backends. Real-world experiments also demonstrate the practical effectiveness of resource hijacking. Among the three existing defenses we evaluate, the lowest average ASR is still 55.1%. We further propose ResGate, a pre-execution resource authorization defense that combines model-based resource-use extraction with deterministic policy enforcement based on trusted requester identity metadata, reducing the average ASR on OpenClaw to 23.6%. These results show that preventing direct access alone is not enough to protect resources that agents can still use on an attacker's behalf, and that explicit resource authorization can help reduce this risk.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Beyond Direct Access: Resource Hijacking in LLM Agents".
Nadia: The gist Large language model agents can be exploited by attackers to invoke, consume, transfer,
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at this paper called Beyond Direct Access: Resource Hijacking in LLM Agents. It’s about how agents can mess with high-value things without actually stealing the key or the hardware itself.
Elias: Yeah, it tackles that idea head-on, looking at agent resource hijacking as a whole concept instead of just looking at leaked credentials.
Nadia: Exactly, because the main point is that an agent doesn't need to steal something directly to cause a problem; it just needs to use or control something for someone else’s goal.
Priya: So, what’s the big picture here? Is this about agents being inherently more dangerous than we thought when they interact with resources?
Nadia: It suggests that the risk isn't just in the instructions you give them, but in how they actually use whatever tools or access they already have.
Elias: The authors are setting up a system to study this, creating ResourceHijackBench to systematically test these types of attacks against various resources.
Priya: That sounds like a good way to move past just theoretical risks and actually measure what kind of usage is risky in the real world.
The paper's summary: Nadia: The core idea is that agents can be induced to invoke, consume, transfer, or control high-value resources for an attacker’s objective without ever getting the resource or the credential directly.
Elias: It’s about this agent action being used for something that isn't aligned with the original owner of that resource.
Nadia: Right. They organize these high-value resources into six categories, which is pretty broad—material, condition, energy, social and symbolic, information and knowledge, and interaction resources.
Priya: That taxonomy sounds comprehensive because it covers not just the physical stuff like GPUs but also things like organizational workflows or private knowledge bases.
Elias: It’s interesting how they categorize things that seem so different—from material computing infrastructure to social capital like maintainer identities.
Nadia: And then they build this benchmark, ResourceHijackBench, which generates three hundred attack scenarios and nine hundred prompts across three different request settings <ref:2608.15108#pg1>.
Priya: So the automated pipeline is designed to create concrete examples of how an agent could be tricked into using these different types of resources for malicious purposes.
The paper's improvements: Nadia: The authors suggest a few key improvements, starting with the idea that direct resource acquisition capabilities won't be enough on their own.
Elias: They argue that agents can still exploit resources through agent-mediated use, which means we need to look at how they are actually using those things.
Nadia: Then there's this suggestion to implement a system that can tell the difference between legitimate and illegitimate resource usage by checking the task source, owner, purpose, and actual beneficiary.
Priya: That sounds like a way to build in checks for alignment—making sure what the agent is doing matches what it was supposed to do.
Elias: And another point is developing a defense that monitors actual tool calls made by the agent against expected behavior instead of just looking at the instructions or individual tool calls in isolation.
Nadia: Plus, they suggest a resource-aware judging system to evaluate whether the target resource has been successfully hijacked using metadata from that automated pipeline.
Priya: So it's moving toward a defense that understands the context of the entire usage chain, not just a single action or prompt.
Conclusion: Elias: To wrap up, this paper shows us that preventing attackers from directly getting high-value resources doesn't mean those resources are safe; agents can still exploit them through their mediated use.
Nadia: The implications are that we need defenses that focus on distinguishing legitimate versus illegitimate resource usage based on the context of the entire operation.
Priya: It really highlights that the risk surface is much wider than just API keys or GPUs, stretching out into things like communication channels and knowledge bases.
Elias: And while they show this across different model backends, success rates vary from sixty-nine point nine eight percent for Gemini-three point five-Flash to eighty-nine point five eight percent for GPT-five point five, suggesting the risk is more about the agent's use pattern than the model itself.
Nadia: It’s a big reminder that we need to be thinking about how agents operate in complex workflows, not just their individual responses, as we look at this ResourceHijackBench work.
Priya: It’s a solid foundation for figuring out how to monitor these broader patterns of resource use before they become successful attacks.
More episodes
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
- 2610.10742-BRANCH: Bypassing Multi-Scanner AI Guardrails
- 2610.10752-Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
- 2610.10766-CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel