When Context Gets Root: Privilege Escalation in LLM Harnesses
cs.CR, cs.SE
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/ByBrawe/opencode-loop
Terminology
Sources
- Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- From Question Answering to Task Completion: A Survey on Agent System and Harness Design
- TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
- AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- Silent Egress: When Implicit Prompt Injection Makes LLM Agents Leak Without a Trace
- Automating Agent Hijacking via Structural Template Injection
- Beyond Resolution Rates: Behavioral Drivers of Coding Agent Success and Failure
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
- When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- Ignore Previous Prompt: Attack Techniques For Language Models
- System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
- Mind the Gap: Action Rebinding Attacks against Android GUI Agents
- Prompt Injection as Role Confusion
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Many-Tier Instruction Hierarchy in LLM Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs