A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
cs.CR, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-08
Comments: 18 pages, 8 figures
Code: https://github.com/HKUDS/OpenHarness
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Dynamic Malicious Skills in Agentic AI
- Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats
- Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance
- ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
- Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
- SMCP: Secure Model Context Protocol
- "Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors
- Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
- From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs