SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
cs.CR, cs.AI
Submitted: 2026-10-07
Updated: 2026-10-07
Code: https://github.com/Stardust457/SLDR
Terminology
Sources
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Defending Against Unforeseen Failure Modes with Latent Adversarial Training
- Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
- Mistral 7B
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
- Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach
- Robustifying Safety-Aligned Large Language Models through Clean Data Curation
- Decoupled Weight Decay Regularization
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Fine-tuning can cripple your foundation model; preserving features may be the solution
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Open Problems in Technical AI Governance
- Representation Noising: A Defence Mechanism Against Harmful Finetuning
- Tamper-Resistant Safeguards for Open-Weight LLMs
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs