Bridging the Gap Between Preference Alignment and Machine Unlearning

arXiv:2504.06659 · cs.LG, cs.AI, cs.CL · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Leveraging Machine Unlearning for Cost-Efficient Preference Alignment".

Jane: The paper was written by Xiaohua Feng, Yuyuan Li, Huwei Ji, Jiaming Zhang, Li Zhang et al. from Zhejiang University and Hangzhou Dianzi University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a fresh one from the arXiv: “Bridging the Gap Between Preference Alignment and Machine Unlearning.” Jane, I’ll be honest, the title alone had me doing a double take.

Jane: Same here, Tom. On the surface, those two ideas sound like opposites. Preference alignment is about teaching a model to be helpful and safe. Machine unlearning is about making it forget specific data. But this paper argues they’re two sides of the same coin.

Tom: Exactly. The old way to align a model is RLHF—reinforcement learning from human feedback. You need tons of examples of *good* responses, which are expensive to collect. And the training is famously unstable and slow.

Jane: Right. And that’s the pain point. The authors say, why do we need all those good examples? What if we just took the *bad* examples—the harmful, biased, or hallucinated responses—and made the model forget them? That’s unlearning.

Tom: And the kicker is, bad examples are everywhere. You get them from user reports, red team tests, even automated checks. They’re cheap. So the whole premise is: can we align a model by subtraction instead of addition?

Jane: But here’s the catch they found early on. It’s not as simple as just unlearning every negative example you can find. The paper shows that unlearning some bad examples actually makes the model *worse* at following human preferences.

Tom: Wait, seriously? You unlearn a harmful response and the model gets *less* aligned?

Jane: That’s what their analysis shows. They built a bi-level optimization framework to measure the impact of unlearning a single sample. And they found the effect can be positive or negative, depending on the sample. It’s not a one-size-fits-all operation.

Tom: So it’s not just “delete the bad stuff.” You have to be picky about *which* bad stuff you delete. That’s a much more nuanced problem than I expected.

Jane: Exactly. And that nuance is the whole reason they built a new framework to do the picking automatically. That’s where the real meat of the paper is, and we’ll get into that next.

Tom: Great, because I want to know how they decide what to forget and what to keep. Let’s dig into the method in a moment.

Summary: Tom: Alright, so we’ve set the stage. The paper is “Bridging the Gap Between Preference Alignment and Machine Unlearning,” and we know the basic idea: use unlearning to align models instead of RLHF. But what’s the actual summary of what they did?

Jane: So they didn’t just try unlearning everything and hope for the best. They first wanted to *quantify* the impact of unlearning a single negative example. They set up this bi-level optimization problem—the inner loop does the unlearning, the outer loop measures the change in alignment performance.

Tom: And that’s where they found the surprising result we teased. The impact is not always positive. They ran experiments on three datasets: one for reducing harmfulness, one for improving usefulness, and one for eliminating hallucinations.

Jane: Right. And they found that whether unlearning helps depends on the *composition* of the sample. They looked at token-level rewards. If a sample has a high proportion of low-reward tokens—meaning most of it is genuinely bad—then unlearning it tends to improve alignment.

Tom: But if it’s only *partially* bad, with a few low-reward tokens mixed in with decent ones, unlearning it can actually hurt. It’s like throwing out the whole recipe because one ingredient is off.

Jane: That’s a great analogy. And that observation led to their main contribution: a framework called U2A, which stands for Unlearning to Align. It’s a bi-level optimization method that automatically selects which negative examples to unlearn and assigns each one a weight.

Tom: So it’s not just a binary “forget or keep.” It’s “how strongly should we forget this one, and should we even bother with that one?”

Jane: Precisely. And the clever part is they use a sparse regularization term to keep the number of selected samples small. They don’t want to unlearn a hundred things if ten will do the job. They even prove that the size of the unlearning set scales with the desired accuracy—so you don’t need to unlearn everything to get close to the optimal result.

Tom: That’s a strong theoretical guarantee. And it means the method is practical, not just a toy. But I’m curious—how well does it actually work in practice? That’s the next segment.

Jane: Good segue, because the results are pretty impressive. Let’s talk about what they actually measured.

Improvements: Tom: So we’ve got the theory. Now, the paper “Bridging the Gap Between Preference Alignment and Machine Unlearning” claims U2A improves things. What did they actually test?

Jane: They took three standard unlearning baselines—GA, GradDiff, and NPO—and plugged them into the U2A framework. Then they compared the improved versions against the originals, and also against standard alignment methods like PPO and DPO.

Tom: And the headline result?

Jane: On the PKU SafeRLHF dataset, which is about reducing harmfulness, the improvements were dramatic. For example, GradDiff alone got a reward value of-four point seven six. With U2A, it jumped to two point six three. That’s a massive swing from negative to positive.

Tom: Whoa. And that’s not just a small bump—that’s going from actively harmful to actually aligned. What about the other metrics?

Jane: They also measured ASR, which is how often the model produces harmful content. On the ASR-answer metric, GradDiff with U2A dropped from fifty-eight point one seven percent to fifty-seven point zero zero percent. Not huge there, but on ASR-useful, it went from twenty-five point four zero percent down to twenty-four point seven five percent. And the model utility, measured by perplexity, actually improved too—from sixty point one one down to fifty-one point two eight.

Tom: So it’s not just better alignment; the model is also more fluent. That’s a win-win. But what about the efficiency claim? Because that was a big part of the pitch.

Jane: That’s the part I love. They set an early stopping condition—stop when the alignment score hits zero point seven four five. PPO needed two hundred eighty-two update rounds. DPO needed one thousand three hundred fifty-three rounds. GradDiff with U2A? Nineteen rounds.

Tom: Nineteen? That’s not even a warm-up.

Jane: Right. And the total time cost for GradDiff with U2A was about seventy seconds, versus over nine hundred seconds for PPO and over three thousand six hundred seconds for DPO. That’s a roughly ninety percent reduction in training time.

Tom: So for a startup or a research lab with limited compute, this is a game-changer. You get better alignment, better utility, and you do it in a fraction of the time. Meng, you’re the engineer here—does that hold up in practice?

Meng: It’s promising, but I’d want to see how it scales to larger models and more diverse datasets. The paper uses 7B and 8B models, which is reasonable, but the real test is whether the selection mechanism stays efficient when you have millions of negative examples to choose from.

Tom: Fair point. But the theoretical complexity analysis suggests it scales linearly with the number of samples, so it’s not obviously a bottleneck. Still, you’re right that real-world deployment is the final exam.

Jane: And that’s the perfect lead-in for our conclusion, where we wrap up what this means for the field.

Conclusion: Tom: Alright, let’s wrap this up. We’ve been deep in “Bridging the Gap Between Preference Alignment and Machine Unlearning,” and I think we can all agree this is a fresh angle on a big problem.

Jane: Absolutely. The core takeaway is that unlearning isn’t just for privacy anymore. It’s a viable path to alignment, and it’s dramatically cheaper than RLHF. The paper shows you don’t need to unlearn everything—just the right things, with the right weights.

Tom: And the numbers back it up. We saw alignment scores go from negative to positive, training time drop by ninety percent, and model utility actually improve. That’s not incremental progress; that’s a paradigm shift in how we think about alignment.

Lu: If I can jump in—what excites me most is the theoretical framing. They’ve given us a way to *measure* the impact of unlearning on alignment. That’s a bridge between two fields that were previously disconnected. Future work can build on this to ask even deeper questions, like how to unlearn in a way that’s robust to distribution shift.

Meng: And from a practical standpoint, this makes alignment accessible to teams that couldn’t afford RLHF. You don’t need a massive preference dataset or a cluster of GPUs for weeks. You need a handful of negative examples and a few hours.

Lalam: And culturally, this matters. It means smaller organizations, educational institutions, and even independent developers can build aligned models. That democratizes AI safety. It’s not just for the big labs anymore.

Tom: That’s a beautiful way to put it, Lalam. So as we say goodbye to this paper, I think the message is clear: alignment doesn’t have to be expensive, and unlearning is more than a privacy tool—it’s a creative lever for making models better.

Jane: Well said. We’ll be back next time with another paper, but for now, thanks for listening, and keep questioning what you think you know about AI.

Tom: See you on the next episode.

Xiaohua Feng, Yuyuan Li, Huwei Ji, Jiaming Zhang, Li Zhang, Tianyu Du, Chaochao Chen

Zhejiang University · Hangzhou Dianzi University

cs.LG, cs.AI, cs.CL

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 17 pages

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 40/100

The gist: Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like Reinforcement Learning with Human Feedback (RLHF) face notable challenges.

Key concepts

Preference Alignment
The process of teaching an AI model to be helpful and safe. Traditionally done via expensive Reinforcement Learning from Human Feedback (RLHF), the paper suggests an alternative approach using unlearning.
Machine Unlearning
The technique of making a trained model 'forget' specific data points, such as harmful or biased examples. The core premise discussed is aligning models by subtraction—removing bad inputs.
RLHF (Reinforcement Learning from Human Feedback)
The traditional, costly method for aligning AI models that requires collecting massive amounts of examples of 'good' responses. The paper argues this process is unstable and slow compared to unlearning.
U2A (Unlearning to Align)
A novel bi-level optimization framework introduced in the paper. It automatically selects which negative examples should be unlearned and assigns them a weight, moving beyond simple deletion for better alignment.

Terminology

Summary

Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like Reinforcement Learning with Human Feedback (RLHF) face notable challenges. These approaches require high-quality datasets of positive preference examples, which are costly to obtain and computationally intensive due to training instability, limiting their use in low-resource scenarios. LLM unlearning technique presents a promising alternative, by directly removing the influence of negative examples. However, current research has primarily focused on empirical validation, lacking systematic quantitative analysis. To bridge this gap, we propose a framework to explore the relationship between PA and LLM unlearning. Specifically, we introduce a bi-level optimization-based method to quantify the impact of unlearning specific negative examples on PA performance. Our analysis reveals that not all negative examples contribute equally to alignment improvement when unlearned, and the effect varies significantly across examples. Building on this insight, we pose a crucial question: how can we optimally select and weight negative examples for unlearning to maximize PA performance? To answer this, we propose a framework called Unlearning to Align (U2A), which leverages bi-level optimization to efficiently select and unlearn examples for optimal PA performance. We validate the proposed method through extensive experiments, with results confirming its effectiveness.

To address the identified challenges, we first develop a special bi-level optimization framework to quantify how unlearning specific negative samples impacts model PA performance. In particular, the inner optimization focuses on unlearning the target sample, while the outer optimization assesses the resulting change in PA performance. After further analysis, we find that not all negative examples contribute to PA improvement, with the degree of impact varying across examples. Meanwhile, the magnitude of the impact is influenced by the unlearning weights. This suggests that indiscriminately applying unlearning to all negative examples fails to achieve optimal PA performance. To address this, we propose a framework called Unlearning to Align (U2A), based on bi-level optimization, to strategically select samples and determine optimal unlearning weights. Further convergence and computational complexity analysis indicate that our proposed method demonstrates good applicability and efficiency in LLMs. This framework bridges the gap between MU and PA, offering a systematic approach to their integration. We summarize the main contributions of this paper as follows:

• We propose a special bi-level optimization framework to measure the impact of unlearning specific samples on PA performance, bridging the gap between MU and PA.

• We find that unlearning all negative examples does not always benefit PA, as their contributions to PA improvement vary and can be adjusted through unlearning weights.

• We propose the U2A framework, leveraging bi-level optimization to select and weight negative examples for unlearning, thereby maximizing PA performance.

• We conduct extensive evaluations on multiple models and real-world datasets, and the experimental results demonstrate the effectiveness of our method.

The mainstream PA approach, RLHF, faces significant challenges in low-resource settings, including (1) reliance on numerous positive preference samples, which are costly to obtain, and (2) instability during training, resulting in high computational and management costs. To address these issues, we propose a MU-based method that reduces dependence on positive samples by mitigating the influence of negative samples to achieve PA. Our method achieves computational efficiency comparable to standard fine-tuning while showing strong potential. We first develop a bi-level optimization framework to evaluate the impact of unlearning individual samples on PA performance. Through this analysis, we observe that negative samples contribute unevenly to PA, with many offering limited benefits. This observation leads to a key question: how can we selectively weight and unlearn negative samples to optimize alignment? To this end, we formally define the problem and introduce U2A, a framework leveraging bi-level optimization to efficiently select and weighted unlearn samples for improved alignment. Experiments demonstrate that U2A significantly enhances alignment efficiency and effectiveness, underscoring its value in resource-constrained scenarios. By linking PA with MU, this work provides a novel perspective on PA for LLMs and suggests new directions for optimizing PA algorithms.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:

Improvements to AI Systems

  1. Implement a Sample-Weighted Unlearning Module for Preference Alignment
  • Action: Instead of applying uniform unlearning to all negative examples, integrate a bi-level optimization module (as described in the U2A framework) that dynamically assigns weights to each negative sample during the unlearning process. This module will use the gradient-based impact analysis (Proposition 4.2) to identify and prioritize samples whose unlearning yields the highest positive impact on preference alignment, while down-weighting or skipping samples with negligible or negative impact.

  • Resulting Capability: The AI system will achieve superior preference alignment (e.g., higher reward scores, lower harmfulness, reduced hallucination) compared to standard unlearning baselines, while using fewer training resources and requiring only negative examples.

  1. Integrate an Adaptive Unlearning Set Selection Mechanism
  • Action: Implement the iterative selection algorithm (Algorithm 1) that builds the unlearning set incrementally. The system will use the marginal gain calculation (Eq. 12) to select the most impactful sample at each iteration, and it will use the early stopping threshold (δ) to automatically determine the optimal number of samples to unlearn, avoiding over-unlearning and preserving model utility.

  • Resulting Capability: The AI system can autonomously determine the minimal and most effective subset of negative data to unlearn for a target alignment level. This leads to faster training convergence (e.g., 90% reduction in training time compared to PPO/DPO in the paper's experiments) and maintains high model utility (low perplexity) on unrelated tasks.

  1. Develop a Low-Resource Alignment Pipeline
  • Action: Replace the computationally expensive RLHF pipeline (which requires costly positive preference data and unstable PPO training) with the U2A-enhanced unlearning pipeline. This involves fine-tuning the base model on a small set of negative examples (which are easier to collect via user reports or red-teaming) using the weighted unlearning loss function from Eq. (9).

  • Resulting Capability: The AI system can be aligned effectively in scenarios with limited data and computational budgets. It will require only negative examples, significantly reducing data collection costs, and its training will be as stable and fast as standard fine-tuning, making it practical for edge devices or rapid iteration cycles.

  1. Add a Quantitative Impact Analysis Tool for Data Auditing
  • Action: Build a diagnostic tool that uses the framework from Section 4.1 to quantify the impact of unlearning any specific data sample on the model's alignment performance. This tool will compute the gradient inner product (Eq. 7) and decompose it into gradient norms and cosine similarity (Eq. 8) to provide actionable insights into why a sample is or isn't beneficial to unlearn.

  • Resulting Capability: The AI system can provide transparency into its own alignment process. Developers can use this tool to audit their negative datasets, identify which samples are truly harmful (low-reward token proportion above a threshold) and which are benign or even beneficial to keep, leading to more informed data curation and better overall model behavior.

Sources

Related papers