Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
cs.CR
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Localizing Model Behavior with Path Patching
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Mistral 7B
- Ignore Previous Prompt: Attack Techniques For Language Models
- Steering Language Models With Activation Engineering
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Qwen3 Technical Report
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs