INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/RebeccaZhang22/intent-as-a-tool
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Reasoning Models Don't Always Say What They Think
- Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Emergent Introspective Awareness in Large Language Models
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Frontier Models are Capable of In-context Scheming
- Large Language Models Often Know When They Are Being Evaluated
- A Survey on Trustworthy LLM Agents: Threats and Countermeasures
- Reasoning Models Struggle to Control their Chains of Thought
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering