The Autonomy Tax: Defense Training Breaks LLM Agents
cs.CR, cs.AI, cs.LG
Submitted: 2026-03-19
Updated: 2026-08-27
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
- Generative AI in Transportation Planning: A Survey
- The Llama 3 Herd of Models
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Deliberative Alignment: Reasoning Enables Safer Language Models
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Mistral 7B
- "Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
- A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
- Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis
- CMOOD: Concept-based Multi-label OOD Detection
- Defenses Against Prompt Attacks Learn Surface Heuristics
- StruQ: Defending Against Prompt Injection with Structured Queries
- Mitigating Hallucinations in Large Language Models via Causal Reasoning
- JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model
- Ignore Previous Prompt: Attack Techniques For Language Models
- Computational complexity of three-dimensional Ising spin glass: Lessons from D-Wave annealer
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs