Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders
cs.CR
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- Sparse Autoencoders are Capable LLM Jailbreak Mitigators
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Toy Models of Superposition
- The Llama 3 Herd of Models
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
- Gemma 2: Improving Open Language Models at a Practical Size
- Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- Qwen2.5 Technical Report
- Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs