Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Yuxi Li, Zhibo Zhang, Kailong Wang, Xingshuo Han, Ling Shi, Haoyu Wang
cs.CR, cs.AI
Submitted: 2026-07-22
Code: https://github.com/tatsu-lab/stanford_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- Benchmarking Ethical and Safety Risks of Healthcare LLMs in China-Toward Systemic Governance under Healthy China 2030
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Evaluating Large Language Models Trained on Code
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
- Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
- Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
- BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
- Backdoor Learning: A Survey
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- The Llama 3 Herd of Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- ONION: A Simple and Effective Defense Against Textual Backdoor Attacks
- Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
- Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Opportunities and Challenges of LLMs in Education: An NLP Perspective
- Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
- Qwen3 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs