What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness
cs.CR
Submitted: 2026-09-01
Updated: 2026-09-02
Code: https://github.com/openai/codex
Project page: https://moonshotai.github.io/kimi-cli/en/customization/agents.html#built-in-subagent-types
License: http://creativecommons.org/licenses/by/4.0/
The gist: Real-world, high-profile AI agent harnesses often rely on vendor-proprietary or opaque designs for context assembly, leaving the sources and underlying logic of assembled context poorly understood
Terminology
Abstract
Real-world, high-profile AI agent harnesses often rely on vendor-proprietary or opaque designs for context assembly, leaving the sources and underlying logic of assembled context poorly understood and the resulting security risks largely unexplored. In this paper, we present the first systematic analysis of context assembly designs in real-world AI agent harnesses. We study and uncover how an agent harness is designed to collect and assemble context from diverse sources, and identify a set of practical attack vectors arising from these designs. Our analysis brings to light two novel categories of attacks in the context assembly of real-world harnesses: (1) MessageRole Context Privilege Escalation (M-CPE), which occurs when attacker-controlled content originating from a low-privileged context is incorporated into a higher-privileged message role. (2) Cross-Scope Context Privilege Escalation (X-CPE), which occurs when attacker-controlled content persists beyond the context in which it was introduced. We performed a systemic security analysis of the CPE attacks against 12 real-world agent harnesses, including Claude Code and Codex. The resulting consequences include full agent compromise, remote code execution, denial of service, and manipulated tool or skill invocations, etc.
Sources
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Prompt Injection Attack to Tool Selection in LLM Agents
- Dynamic Malicious Skills in Agentic AI
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Lessons from Defending Gemini Against Indirect Prompt Injections
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs