Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
cs.CR, cs.AI
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Knowledge Unlearning for Mitigating Privacy Risks in Language Models
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- From Cognition to Computation: A Comparative Review of Human Attention and Transformer Architectures
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs