CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
cs.CR, cs.AI
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- StruQ: Defending Against Prompt Injection with Structured Queries
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection
- ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
- CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing
- Steering Instruction Hierarchies at Inference Time
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Steering Language Models With Activation Engineering
- Steering Llama 2 via Contrastive Activation Addition
- Refusal in Language Models Is Mediated by a Single Direction
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen3 Technical Report
- InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs