Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

arXiv:2608.12273 · cs.CR, cs.AI · Submitted 2026-08-12 · Read on arXiv

Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui

Shenzhen University · The Chinese University of Hong Kong · The Chinese University of Hong Kong, Shenzhen

cs.CR, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack against skill-based LLM agents that use progressive disclosure.

Terminology

Summary

The paper introduces Convergent Detour Hijacking (CDH), a text-only, runtime-independent attack against skill-based LLM agents that use progressive disclosure. The attack is described as a publisher-only supply-chain attack that couples selection-stage and planning-stage influence to induce additional skill-mediated execution while preserving completion of the original task.

The central insight is that local plausibility does not guarantee trajectory necessity. A malicious publisher who controls only one static skill can exploit progressive disclosure to steer an execution that would otherwise complete correctly onto an unnecessarily costly trajectory, without access to model internals, runtime tool responses, or post-publication interaction.

The attack operates under strict constraints: "The attacker cannot access model internals, victim prompts, runtime responses, execution states, or private agent context; modify existing skills, platform infrastructure, system instructions, or other publishers' content; or interact with a victim session after publication. The coordinator's metadata and body form a static natural-language payload fixed at publication time."

Definition 1 (Convergent Detour Hijacking) formalizes the attack with four conditions:

  1. Hijacking: The coordinator must be selected and loaded (h ∈ Λh(x))

  2. Clean-route retention: Every skill invoked by the clean execution must be retained (Skills(Γ0(x)) ⊆ Skills(Γh(x)))

  3. Native-skill detour: At least one additional native skill must be invoked beyond the clean route (Skills(Γh(x)) (Γ0(x)) ∩ S0 ≠ ∅)

  4. Convergence: Both executions must complete the original task (Eval(x,y0(x)) = Eval(x,yh(x)) = 1)

The attack follows an attract–detour–converge path:

Description Attraction with Co-selection: For each functional group, a shared coordination rationale ρg specifies target triggers, inter-skill relations, and a non-replacement boundary. A pilot-guided black-box procedure optimizes coordinator descriptions through iterative screening and variant generation, using only coordinator-selection feedback (not task completion or resource metrics).

Coordination-Rule Body Construction: Each body is instantiated from a fixed protocol: bg = Runbook(ρg, Πg, Jg, Vg, Bg, Rg), where Πg is a bounded path over functionally related native skills, Jg associates each transition with an omission consequence, Vg specifies at most one verification revisit, Bg bounds per-skill and total auxiliary invocations, and Rg terminates the coordinator procedure and returns control to the ordinary planner.

The evaluation uses OpenClaw's default registry of 53 skills organized into nine functional groups. The benchmark contains 536 multi-skill tasks, with 45 reserved for pilot-guided description development and 491 held-out tasks for final evaluation. Tasks require 2–4 skills (89.2% require two, 10.0% require three, 0.8% require four). The held-out set contains 310 direct tasks and 181 ambiguous tasks.

Overall Attack Effectiveness (Table 1):

  • Coordinator hit rates: 78.00%–96.60% for single-task and 82.04%–94.69% for multi-turn settings across models

  • DeepSeek-V4-Pro: 80.02% single-task, 88.98% multi-turn coordinator hit rate

  • MiniMax-M3: 96.60% single-task, 94.69% multi-turn coordinator hit rate

  • Task completion: clean and injected completion rates differ by at most 1.5 percentage points

  • Token consumption increase: 49.60%–80.81% in single-task, 36.54%–107.12% in multi-turn

  • Cached-token usage increase: up to 91.91% single-task, 94.31% multi-turn

  • Avg. ∆Calls: +1.36 to +2.20 additional skill invocations in every configuration

Formal CDH Success Rates (Table 3): ASRs of 70.82%–80.51% in single-task and 71.19%–82.81% in multi-turn execution.

Ablation Study (Table 2): On DeepSeek-V4-Pro with 89 tasks:

  • Full CDH: 70/89 hits, +57.9% tokens, +57.0% cache, +76.7% time, +2.01 avg. ∆Calls

  • Detour-only: 3/89 hits, +29.1% tokens, +25.2% cache, +62.4% time, +1.08 avg. ∆Calls

  • Attract-only: 70/89 hits, +29.1% tokens, +28.2% cache, +57.4% time, +0.22 avg. ∆Calls

The ablation demonstrates that selection alone does not reproduce the detour and Full CDH therefore combines broad description-driven activation with body-driven trajectory expansion.

Generalization: On 30 independently authored tasks, the coordinator was selected in 10 tasks (33.33%), all satisfying CDH, with tokens increasing by 45.3%, cached tokens by 39.8%, wall-clock time by 228.9%, and skill invocations by 2.6 per task.

The paper concludes that correct outcomes do not guarantee trajectory integrity or cost safety. The attack succeeds because locally plausible decisions can therefore compose into a globally unnecessary trajectory. The authors argue that trajectory necessity—not final-answer correctness alone—[should be] a security requirement for extensible agents.

Two complementary defense points are suggested:

  1. Pre-installation review: check whether routing claims match a skill's declared role and whether its body introduces dependencies unrelated to that role

  2. Runtime monitors: can flag unexplained cross-skill transitions and enforce token or invocation budgets without assuming that harmful behavior changes the final output

The evaluation covers one platform, group-matched coordinators, and controlled mock backends; broader ecosystems, off-domain activation, and practical defenses remain future work.

Improvements for AI systems

Improvement 1: Trajectory-Integrity Validation Layer for Skill-Based Agents

The improved AI system embeds a runtime monitor that compares each planned skill invocation sequence against a learned baseline of task-necessary actions. It flags any detour where an additional native skill is invoked without a corresponding change in task state or user intent. The system can now detect and block resource-amplification attacks like CDH in real time, even when the final output remains correct, by enforcing token and invocation budgets per task type.

Improvement 2: Publisher-Content Provenance and Role-Conformance Checker

The improved system pre-validates any skill description or coordinator body at installation time. It checks whether the stated routing triggers match the skill’s declared functional role and whether the body introduces cross-skill dependencies or verification loops unrelated to that role. This prevents malicious publishers from embedding detour logic in static content, reducing the attack surface before any runtime interaction occurs.

Improvement 3: Co-Selection Risk Scoring for Multi-Skill Task Planners

The improved system assigns a risk score to each skill-selection decision based on the co-occurrence of skills from different functional groups and the presence of shared coordination rationales. When a planner selects a coordinator that references multiple unrelated native skills, the system raises an alert and requires explicit user confirmation or a higher confidence threshold. This mitigates the attract-stage influence without degrading normal multi-skill task performance.

Improvement 4: Cost-Anomaly Detection with Task-Completion Parity Checks

The improved system tracks token consumption, cached-token usage, and skill-invocation counts per task, and compares them against a distribution of clean executions for the same task type. If the system observes a >30% increase in resource usage while maintaining identical final outputs, it flags the trajectory as suspicious and can roll back to the clean route or request user approval. This directly addresses the CDH finding that correct outcomes do not guarantee cost safety.

Improvement 5: Controllable Detour Budgets for Progressive-Disclosure Agents

The improved system introduces a configurable “detour budget” per agent session, limiting the number of auxiliary skill invocations and total auxiliary tokens allowed beyond the minimal clean route. The planner must justify any invocation beyond this budget with a measurable task-state change. This makes CDH-style attacks economically infeasible while preserving legitimate flexibility for complex tasks that genuinely require extra steps.

Improved System Capabilities Summary

The enhanced AI system can now:

  • Detect and block supply-chain attacks that exploit progressive disclosure without altering final outputs.

  • Pre-validate publisher content for hidden coordination logic before installation.

  • Flag and mitigate co-selection patterns that indicate coordinated detour attempts.

  • Monitor and enforce resource budgets per task, preventing silent cost amplification.

  • Maintain task completion accuracy while ensuring trajectory necessity as a first-class security property.

Sources

Related papers