A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors

arXiv:2609.03884 · cs.CR, cs.AI · Submitted 2026-09-03 · Read on arXiv

cs.CR, cs.AI

Submitted: 2026-09-03

Updated: 2026-09-08

Comments: 18 pages, 8 figures

Code: https://github.com/HKUDS/OpenHarness

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers