LLM Unlearning Evaluation with TRIAGE
cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/Dan-A2/Concept-Unlearning
Terminology
Sources
- To Each (Textual Sequence) Its Own: Improving Memorized-Data Unlearning in Large Language Models
- Discovering Latent Knowledge in Language Models Without Supervision
- Do Unlearning Methods Remove Information from Language Model Weights?
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- Studying Large Language Model Generalization with Influence Functions
- Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
- Measuring Massive Multitask Language Understanding
- LoRA: Low-Rank Adaptation of Large Language Models
- BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
- On Effects of Steering Latent Representation for Large Language Model Unlearning
- Improving LLM Unlearning Robustness via Random Perturbations
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
- Eight Methods to Evaluate Robust Unlearning in LLMs
- TOFU: A Task of Fictitious Unlearning for LLMs
- Scalable Extraction of Training Data from (Production) Language Models
- UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
- Align-then-Unlearn: Embedding Alignment for LLM Unlearning
- Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
- Is your algorithm unlearning or untraining?
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks