RADNPO: Reference-free Adaptive Negative Preference Optimization for LLM Unlearning
cs.CR
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- KTO: Model Alignment as Prospect Theoretic Optimization
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
- Federated Unlearning: How to Efficiently Erase a Client in FL?
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- TOFU: A Task of Fictitious Unlearning for LLMs
- Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
- In-Context Unlearning: Language Models as Few Shot Unlearners
- Proximal Policy Optimization Algorithms
- MUSE: Machine Unlearning Six-Way Evaluation for Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
- Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
- CALIBURN: Self-Calibrated LLM Unlearning Alignment
- Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs