Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning
cs.CR, cs.AI, cs.IR, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Defending Against Unforeseen Failure Modes with Latent Adversarial Training
- Can Editing LLMs Inject Harm?
- LoRA: Low-Rank Adaptation of Large Language Models
- Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
- Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
- Evaluating Defences against Unsafe Feedback in RLHF
- Representation Noising: A Defence Mechanism Against Harmful Finetuning
- Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
- Locking Down the Finetuned LLMs Safety
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs