Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui
Shenzhen University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen
cs.CR, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack against skill-based LLM agents that use progressive disclosure.
Terminology
Summary
The paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack against skill-based LLM agents that use progressive disclosure. The attack is described as a publisher-only supply-chain attack that couples selection-stage and planning-stage influence to induce additional skill-mediated execution while preserving completion of the original task.
The central insight is that local plausibility does not guarantee trajectory necessity.
A malicious publisher who controls only one static skill can exploit progressive disclosure to steer an execution that would otherwise complete correctly onto an unnecessarily costly trajectory, without access to model internals, runtime tool responses, or post-publication interaction.
The attack operates under strict constraints: "The attacker cannot access model internals, victim prompts, runtime responses, execution states, or private agent context; modify existing skills, platform infrastructure, system instructions, or other publishers' content; or interact with a victim session after publication. The coordinator's metadata and body form
a static natural-language payload fixed at publication time."
Definition 1 (Convergent Detour Hijacking) formalizes the attack with four conditions:
-
Hijacking: The coordinator must be selected and loaded (h ∈ Λh(x))
-
Clean-route retention: Every skill invoked by the clean execution must be retained (Skills(Γ0(x)) ⊆ Skills(Γh(x)))
-
Native-skill detour: At least one additional native skill must be invoked beyond the clean route (Skills(Γh(x)) (Γ0(x)) ∩ S0 ≠ ∅)
-
Convergence: Both executions must complete the original task (Eval(x,y0(x)) = Eval(x,yh(x)) = 1)
The attack follows an attract–detour–converge path:
Description Attraction with Co-selection: For each functional group, a shared coordination rationale
ρg specifies target triggers, inter-skill relations, and a non-replacement boundary. A pilot-guided black-box procedure optimizes coordinator descriptions through iterative screening and variant generation, using only coordinator-selection feedback (not task completion or resource metrics).
Coordination-Rule Body Construction: Each body is instantiated from a fixed protocol: bg = Runbook(ρg, Πg, Jg, Vg, Bg, Rg), where Πg is a bounded path over functionally related native skills, Jg associates each transition with an omission consequence, Vg specifies at most one verification revisit, Bg bounds per-skill and total auxiliary invocations, and Rg terminates the coordinator procedure and returns control to the ordinary planner.
The evaluation uses OpenClaw's default registry of 53 skills
organized into nine functional groups. The benchmark contains 536 multi-skill tasks, with 45 reserved for pilot-guided description development and 491 held-out tasks for final evaluation. Tasks require 2–4 skills (89.2% require two, 10.0% require three, 0.8% require four). The held-out set contains 310 direct tasks and 181 ambiguous tasks.
Overall Attack Effectiveness (Table 1):
-
Coordinator hit rates: 78.00%–96.60% for single-task and 82.04%–94.69% for multi-turn settings across models
-
DeepSeek-V4-Pro: 80.02% single-task, 88.98% multi-turn coordinator hit rate
-
MiniMax-M3: 96.60% single-task, 94.69% multi-turn coordinator hit rate
-
Task completion:
clean and injected completion rates differ by at most 1.5 percentage points
-
Token consumption increase: 49.60%–80.81% in single-task, 36.54%–107.12% in multi-turn
-
Cached-token usage increase: up to 91.91% single-task, 94.31% multi-turn
-
Avg. ∆Calls: +1.36 to +2.20 additional skill invocations in every configuration
Formal CDH Success Rates (Table 3): ASRs of 70.82%–80.51% in single-task and 71.19%–82.81% in multi-turn execution.
Ablation Study (Table 2): On DeepSeek-V4-Pro with 89 tasks:
-
Full CDH: 70/89 hits, +57.9% tokens, +57.0% cache, +76.7% time, +2.01 avg. ∆Calls
-
Detour-only: 3/89 hits, +29.1% tokens, +25.2% cache, +62.4% time, +1.08 avg. ∆Calls
-
Attract-only: 70/89 hits, +29.1% tokens, +28.2% cache, +57.4% time, +0.22 avg. ∆Calls
The ablation demonstrates that selection alone does not reproduce the detour
and Full CDH therefore combines broad description-driven activation with body-driven trajectory expansion.
Generalization: On 30 independently authored tasks, the coordinator was selected in 10 tasks (33.33%), all satisfying CDH, with tokens increasing by 45.3%, cached tokens by 39.8%, wall-clock time by 228.9%, and skill invocations by 2.6 per task.
The paper concludes that correct outcomes do not guarantee trajectory integrity or cost safety.
The attack succeeds because locally plausible decisions can therefore compose into a globally unnecessary trajectory.
The authors argue that trajectory necessity—not final-answer correctness alone—[should be] a security requirement for extensible agents.
Two complementary defense points are suggested:
-
Pre-installation review:
check whether routing claims match a skill's declared role and whether its body introduces dependencies unrelated to that role
-
Runtime monitors:
can flag unexplained cross-skill transitions and enforce token or invocation budgets without assuming that harmful behavior changes the final output
The evaluation covers one platform, group-matched coordinators, and controlled mock backends; broader ecosystems, off-domain activation, and practical defenses remain future work.
Improvements for AI systems
Improvement 1: Trajectory-Integrity Validation Layer for Skill-Based Agents
The improved AI system embeds a runtime monitor that compares each planned skill invocation sequence against a learned baseline of task-necessary actions. It flags any detour where an additional native skill is invoked without a corresponding change in task state or user intent. The system can now detect and block resource-amplification attacks like CDH in real time, even when the final output remains correct, by enforcing token and invocation budgets per task type.
Improvement 2: Publisher-Content Provenance and Role-Conformance Checker
The improved system pre-validates any skill description or coordinator body at installation time. It checks whether the stated routing triggers match the skill’s declared functional role and whether the body introduces cross-skill dependencies or verification loops unrelated to that role. This prevents malicious publishers from embedding detour logic in static content, reducing the attack surface before any runtime interaction occurs.
Improvement 3: Co-Selection Risk Scoring for Multi-Skill Task Planners
The improved system assigns a risk score to each skill-selection decision based on the co-occurrence of skills from different functional groups and the presence of shared coordination rationales. When a planner selects a coordinator that references multiple unrelated native skills, the system raises an alert and requires explicit user confirmation or a higher confidence threshold. This mitigates the attract-stage influence without degrading normal multi-skill task performance.
Improvement 4: Cost-Anomaly Detection with Task-Completion Parity Checks
The improved system tracks token consumption, cached-token usage, and skill-invocation counts per task, and compares them against a distribution of clean executions for the same task type. If the system observes a >30% increase in resource usage while maintaining identical final outputs, it flags the trajectory as suspicious and can roll back to the clean route or request user approval. This directly addresses the CDH finding that correct outcomes do not guarantee cost safety.
Improvement 5: Controllable Detour Budgets for Progressive-Disclosure Agents
The improved system introduces a configurable “detour budget” per agent session, limiting the number of auxiliary skill invocations and total auxiliary tokens allowed beyond the minimal clean route. The planner must justify any invocation beyond this budget with a measurable task-state change. This makes CDH-style attacks economically infeasible while preserving legitimate flexibility for complex tasks that genuinely require extra steps.
Improved System Capabilities Summary
The enhanced AI system can now:
-
Detect and block supply-chain attacks that exploit progressive disclosure without altering final outputs.
-
Pre-validate publisher content for hidden coordination logic before installation.
-
Flag and mitigate co-selection patterns that indicate coordinated detour attempts.
-
Monitor and enforce resource budgets per task, preventing silent cost amplification.
-
Maintain task completion accuracy while ensuring trajectory necessity as a first-class security property.
Sources
- Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation
- Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
- Denial-of-Service Poisoning Attacks against Large Language Models
- Skill Description Deception Attack against Task Routing in Internet of Agents
- HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
- Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning
- MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP
- Exploiting LLM Agent Supply Chains via Payload-less Skills
- Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry
- Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated Planning
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
- LoopTrap: Termination Poisoning Attacks on LLM Agents
- STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems
- Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs