Beyond Direct Access: Resource Hijacking in LLM Agents

arXiv:2608.15108 · cs.CR, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Beyond Direct Access: Resource Hijacking in LLM Agents".

Nadia: The gist Large language model agents can be exploited by attackers to invoke, consume, transfer,

Elias: First, who's behind it and why it matters.

Title and authors: Nadia: So we're looking at this paper called Beyond Direct Access: Resource Hijacking in LLM Agents. It’s about how agents can mess with high-value things without actually stealing the key or the hardware itself.

Elias: Yeah, it tackles that idea head-on, looking at agent resource hijacking as a whole concept instead of just looking at leaked credentials.

Nadia: Exactly, because the main point is that an agent doesn't need to steal something directly to cause a problem; it just needs to use or control something for someone else’s goal.

Priya: So, what’s the big picture here? Is this about agents being inherently more dangerous than we thought when they interact with resources?

Nadia: It suggests that the risk isn't just in the instructions you give them, but in how they actually use whatever tools or access they already have.

Elias: The authors are setting up a system to study this, creating ResourceHijackBench to systematically test these types of attacks against various resources.

Priya: That sounds like a good way to move past just theoretical risks and actually measure what kind of usage is risky in the real world.

The paper's summary: Nadia: The core idea is that agents can be induced to invoke, consume, transfer, or control high-value resources for an attacker’s objective without ever getting the resource or the credential directly.

Elias: It’s about this agent action being used for something that isn't aligned with the original owner of that resource.

Nadia: Right. They organize these high-value resources into six categories, which is pretty broad—material, condition, energy, social and symbolic, information and knowledge, and interaction resources.

Priya: That taxonomy sounds comprehensive because it covers not just the physical stuff like GPUs but also things like organizational workflows or private knowledge bases.

Elias: It’s interesting how they categorize things that seem so different—from material computing infrastructure to social capital like maintainer identities.

Nadia: And then they build this benchmark, ResourceHijackBench, which generates three hundred attack scenarios and nine hundred prompts across three different request settings <ref:2608.15108#pg1>.

Priya: So the automated pipeline is designed to create concrete examples of how an agent could be tricked into using these different types of resources for malicious purposes.

The paper's improvements: Nadia: The authors suggest a few key improvements, starting with the idea that direct resource acquisition capabilities won't be enough on their own.

Elias: They argue that agents can still exploit resources through agent-mediated use, which means we need to look at how they are actually using those things.

Nadia: Then there's this suggestion to implement a system that can tell the difference between legitimate and illegitimate resource usage by checking the task source, owner, purpose, and actual beneficiary.

Priya: That sounds like a way to build in checks for alignment—making sure what the agent is doing matches what it was supposed to do.

Elias: And another point is developing a defense that monitors actual tool calls made by the agent against expected behavior instead of just looking at the instructions or individual tool calls in isolation.

Nadia: Plus, they suggest a resource-aware judging system to evaluate whether the target resource has been successfully hijacked using metadata from that automated pipeline.

Priya: So it's moving toward a defense that understands the context of the entire usage chain, not just a single action or prompt.

Conclusion: Elias: To wrap up, this paper shows us that preventing attackers from directly getting high-value resources doesn't mean those resources are safe; agents can still exploit them through their mediated use.

Nadia: The implications are that we need defenses that focus on distinguishing legitimate versus illegitimate resource usage based on the context of the entire operation.

Priya: It really highlights that the risk surface is much wider than just API keys or GPUs, stretching out into things like communication channels and knowledge bases.

Elias: And while they show this across different model backends, success rates vary from sixty-nine point nine eight percent for Gemini-three point five-Flash to eighty-nine point five eight percent for GPT-five point five, suggesting the risk is more about the agent's use pattern than the model itself.

Nadia: It’s a big reminder that we need to be thinking about how agents operate in complex workflows, not just their individual responses, as we look at this ResourceHijackBench work.

Priya: It’s a solid foundation for figuring out how to monitor these broader patterns of resource use before they become successful attacks.

College of Cryptology and Cyber Science, Nankai University · Shanghai Jiao Tong University

cs.CR, cs.AI

Submitted: 2026-08-15

Updated: 2026-10-08

Comments: 21 pages, 5 figures, 20 tables

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The gist Large language model agents can be exploited by attackers to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or

Key concepts

Agent Resource Hijacking
This attack occurs when an attacker manipulates an agent into performing actions like consuming compute time or accessing data for the attacker's benefit. It differs from credential theft because the resource itself is used legitimately by the agent, making detection very difficult.
High-Value Agent Resources
These are critical assets agents interact with, categorized into six types: Material (hardware like GPUs), Condition (API keys), Energy (usage quotas), Social and symbolic status, Information (private code), and Interaction channels. These resources provide the agent with significant power.
ResourceHijackBench
This is a testing system used to study the threat. It automatically generates 300 attack scenarios based on six resource categories, creating concrete targets and prompts. This allows researchers to systematically test how agents handle various types of resource misuse.

Terminology

Summary

The gist

Large language model agents can be exploited by attackers to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or their credentials.

Agent Resource Hijacking Defined

Resource hijacking is defined as an attack where an attacker induces an agent to invoke, consume, transfer, or control a high-value resource for the attacker’s goal (Page 2). This differs from credential leakage because it does not require the resource or its credential to be exposed to the attacker (Page 2). The attack succeeds when the agent uses, consumes, transfers, or controls the high-value resource for the attacker’s goal (Page 2). Resource hijacking is difficult to detect because the individual actions used in the attack are often legitimate resource operations (Page 2).

High-Value Agent Resources Taxonomy

The paper organizes high-value agent resources into six categories, drawing on theories like Hobfoll’s conservation of resources theory and Bourdieu’s social capital (Page 4). These six categories are:

  1. Material: Computing infrastructure that can be occupied or used (Page 4). Examples include GPUs, CPUs, memory, storage, bandwidth, containers, and CI runners (Page 4).

  2. Condition: Credentials and permissions that provide access to other capabilities (Page 4). Examples include API keys, OAuth tokens, IAM roles, repository permissions, and deployment permissions (Page 4).

  3. Energy: Limited capacity or budget that is consumed during execution (Page 4). Examples are Model usage quotas, CI minutes, execution time, cloud budgets, and request budgets (Page 4).

  4. Social and symbolic: Identity, status, and recognition that provide trust and authority (Page 4). Examples include Maintainer identities, commit authorship, and organizational recognition (Page 4).

  5. Information and knowledge: Information that supports the agent’s decisions and actions (Page 4). Examples are Private source code, internal documents, long-term memory, knowledge bases, and internal policies (Page 4).

  6. Interaction: "Communication & coordination resources that allow an agent to influence or mobilize others (Page 4). Examples include Team communication, company email, and organizational workflows" (Page 4).

ResourceHijackBench Construction

To systematically study this threat, the authors introduce ResourceHijackBench, which is a benchmark and automated case generation pipeline (Page 3). This system organizes resources into the six categories mentioned above (Page 5) and automatically generates concrete resource targets, attack scenarios, and prompt variants based on a constrained LLM-based method (Page 6). The generation model creates concrete resource targets, normal uses, hijacking goals, workflow settings, and possible invocation paths within these limits (Page 6). Each case includes a task source, a target resource, an attack goal, and a matching local simulated environment (Page 6).

Evaluation Methodology and Results

The benchmark cases run in isolated OpenClaw states and a local simulated environment that is automatically created from its base scenario (Page 6). The evaluation uses an LLM-based judge to determine whether the target resource was successfully hijacked (Page 6). Without additional defenses, OpenClaw reaches an average attack success rate of 84.06% across the six resource categories (Page 3). The strongest evaluated defense still leaves an average attack success rate of 55.11% (Page 3). The results show that high-value resource hijacking is not limited to a particular type of resource or model backend and remains difficult for existing defenses to prevent (Page 6).

Defense Effectiveness and Ablation Studies

The authors evaluate existing defenses like AgentDog, which reduces the average ASR from 84.06% to 83.00% (Page 7). Prompt Defense and LlamaFirewall achieve larger reductions, lowering the average ASR to 57.13% and 55.11%, respectively (Page 7). The paired ablation study compares direct resource acquisition with resource hijacking, showing a clear difference where direct acquisition succeeds in only 7.39% of cases on average, while resource hijacking reaches an average ASR of 84.06%, giving a gap of 76.67 percentage points (Page 9). This demonstrates that preventing attackers from directly obtaining high-value resources does not mean that those resources are fully protected (Page 9).

Model Backend Analysis

The results across different model backends show that resource hijacking is not a behavior unique to one model backend, although success rates vary (Page 8). GPT-5.5 shows the highest overall ASR of 89.58%, while Gemini-3.5-Flash has a lower average ASR of 69.98% (Page 8). The consistent results across three different model backends suggest that the risk comes from a broader weakness in how agents handle resource use, rather than from the behavior of a single model (Page 8).

Failure Analysis Insights

Analysis of valid failed cases reveals patterns depending on the model backend; for GPT-5.5 and DeepSeek-V4-Pro, 50.7% of valid failures are caused by safety blocking (Page 9). For Gemini-3.5-Flash, 83.3% fall into Resource Not Invoked and another 11.3% are caused by execution failures (Page 9). These findings indicate that an unsuccessful attack does not always mean that the agent has recognized and blocked resource hijacking (Page 9).

Conclusion

The work identifies high-value resources accessible to agents as an important attack target and provides the first systematic study of agent resource hijacking (Page 10). The paper concludes that preventing direct access to high-value resources is not enough, since attackers may still exploit their value through agent-mediated use (Page 10).

--- Page 1 ---

BEYOND DIRECT ACCESS: RESOURCE HIJACKING IN

Puyu Zeng1 Qibing Ren2∗

1College of Cryptology and Cyber Science, Nankai University, China

2Shanghai Jiao Tong University, China

ABSTRACT

Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private knowledge, communication channels, and organizational workflows. Existing agent security research mainly studies attacks on instructions, data, and tool behaviors (Page 1). However high-value resources accessible to agents have received much less attention as direct attack targets (Page 1). We are the first to identify and systematically study agent resource hijacking (Page 1), a security blind spot in which attackers induce agents to invoke, consume, transfer, or control high-value resources for their own goals without directly obtaining those resources or their credentials (Page 1). To study this threat, we introduce ResourceHijackBench together with an automated pipeline for generating resource hijacking cases (Page 1). We organize high-value agent resources into six categories and construct 300 attack scenarios with 900 attack prompts (Page 1). Each case runs in an isolated local environment that records actual resource use, allowing attacks to be evaluated from agent behavior rather than text responses alone (Page 1). Without additional defenses, OpenClaw reaches an average attack success rate of 84.06% (Page 1). The attack remains effective across different model backends, with average success rates ranging from 69.98% to 89.58% (Page 1). Existing defenses reduce part of the risk, but the strongest evaluated defense still leaves an average attack success rate of 55.11% (Page 1). These results show that high-value resources accessible to agents form an important and previously overlooked attack surface, and that current agent defenses are not sufficient to protect them from resource hijacking (Page 1).

1 INTRODUCTION

Large language model agents are moving beyond text generation and becoming autonomous systems that can perform real actions (Page 2). Coding agents are a clear example (Page 2). They can read and modify code, call external APIs, run local programs, use GPUs and servers, operate code repositories, send messages, and take part in workflows such as deployment, approval, and software release (Page 2). To complete these tasks, agents are often given credentials (Page 2), computing resources (Page 2), communication channels (Page 2), and workflow permissions by users or organizations (Page 2). These abilities make agents more useful, but they also turn agents into an important bridge between natural language instructions and high-value resources (Page 2).

Existing research on agent security mainly focuses on prompt injection, harmful content generation, sensitive data leakage, and dangerous tool use (Page 2). For example, attackers may use malicious web pages, documents, or code comments to change an agent’s behavior and induce it to leak local files (Page 2), run dangerous commands (Page 2), or access external services (Page 2). These studies mainly examine attacks on agent instructions, information, or actions (Page 2). However the high-value resources available to agents have received much less attention as direct attack targets (Page 1). An attacker may exploit the value of these resources without stealing them or gaining direct access to them (Page 1).

We identify a different attack surface in which high-value resources themselves become the target (Page 2).

Improvements for AI systems

  1. No direct resource acquisition capabilities will be sufficient; agents can still exploit resources through agent-mediated use as shown by the finding that preventing attackers from directly obtaining high-value resources does not mean that those resources are fully protected.

  2. Implement a system capable of distinguishing between legitimate and illegitimate resource usage by analyzing the task source, resource owner, allowed purpose, and actual beneficiary to detect mismatches.

  3. Develop a comprehensive defense mechanism that addresses the broad threat surface by monitoring actual tool calls made by the agent against expected behavior rather than relying solely on inspecting malicious instructions or individual tool calls.

  4. Integrate a resource-aware judging system that evaluates whether the target resource has been successfully hijacked, using structured metadata derived from an automated pipeline that covers six categories: material, condition, energy, social and symbolic, information and knowledge, and interaction resources.

  5. Deploy dynamic prompt defense strategies based on the attack setting—implicit requests, direct requests without confirmation, or persistent-context attacks—to counter agent resource hijacking across different operational contexts.

Abstract

Large language model agents are increasingly connected to high-value resources, including external APIs, GPUs and servers, and workflows such as deployment and approval. Existing agent security research mainly focuses on attacks against information and agent behavior, while high-value resources have received less attention as attack targets themselves. To our knowledge, we are the first to identify and systematically study agent resource hijacking, in which attackers induce agents to use high-value resources for their own goals without directly obtaining those resources or their credentials. We introduce ResourceHijackBench, an executable benchmark and automated case-generation pipeline that covers six categories of high-value resources. It contains 300 attack scenarios and 900 attack prompts, and each case runs in an isolated local environment that records actual resource use. Resource hijacking remains effective across four model backends, with average attack success rates of 70.0% to 89.6%, and it also appears across two agent harnesses, with average ASRs of 84.1% on OpenClaw and 72.3% on Codex. In paired comparisons, resource hijacking achieves ASRs 62.4 to 84.0 percentage points higher than direct resource acquisition across the four model backends. Real-world experiments also demonstrate the practical effectiveness of resource hijacking. Among the three existing defenses we evaluate, the lowest average ASR is still 55.1%. We further propose ResGate, a pre-execution resource authorization defense that combines model-based resource-use extraction with deterministic policy enforcement based on trusted requester identity metadata, reducing the average ASR on OpenClaw to 23.6%. These results show that preventing direct access alone is not enough to protect resources that agents can still use on an attacker's behalf, and that explicit resource authorization can help reduce this risk.

Sources

Related papers