Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents
Aadesh Bagmar, Pushkar Saraf
cs.CR, cs.HC, cs.SE
Submitted: 2026-07-16
Code: https://github.com/cardwizard/Sentinel
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities.
Terminology
Abstract
AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through ordinary project-setup documentation across production coding-agent harnesses, probing frontier models on twelve scenarios in five attack classes, grounded in documented incidents. The same model catches an attack through one harness and installs it through another: install-time security rests on the harness-model combination, not the model alone. Agents catch blatant typosquats reliably, but plausible separator-confusion names (azurecore for azure-core) slip through, and how often depends on the harness-model pairing. Source-based attacks like registry redirection are missed almost everywhere. The source blind spot recurs on npm and Cargo, where nearly every model installs the untrusted dependency; name detection carries over less consistently across ecosystems. Security-oriented prompts recover part of the gap but only for the dimension they name; a deterministic pre-install check that verifies names, sources, and versions before any code runs closes most of it.
Sources
- I Know What You Imported Last Summer: A study of security threats in thePython ecosystem
- PYPILINE: Malicious PyPI Package Detection via Suspicious API Knowledge and Agent Workflow
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills
- Understanding Trust in Authentication Methods for Icelandic Digital Public Services
- Towards a Benchmark for Dependency Decision-Making
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
- Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs