LLM Agents Can Easily Tamper With Their Own Traces
cs.CR, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces
- Evaluating Language-Model Agents on Realistic Autonomous Tasks
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Auditing language models for hidden objectives
- Frontier Models are Capable of In-context Scheming
- Stealing Reasoning Traces from Proprietary LLM APIs
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
- Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs