You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
cs.CR
Submitted: 2026-05-06
Updated: 2026-08-27
Code: https://github.com/AI4LIFE-GROUP/med-safety
Terminology
Sources
- Concrete Problems in AI Safety
- SecureBreak -- A dataset towards safe and secure models
- XBreaking: Understanding how LLMs security alignment can be broken
- LoRA as Oracle
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Unveiling the Basin-Like Loss Landscape in Large Language Models
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- Training Verifiers to Solve Math Word Problems
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- The Aloe Family Recipe for Open and Specialized Healthcare LLMs
- The Llama 3 Herd of Models
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- Measuring Mathematical Problem Solving With the MATH Dataset
- LoRA: Low-Rank Adaptation of Large Language Models
- UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs