Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

arXiv:2509.03647 · cs.CL, cs.AI, cs.LG · Submitted 2025-09-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Breaking the Mirror".

Jane: Large language models (LLMs) used as evaluators suffer from self-preference bias, where they disproportionately favor their own outputs, which undermines fairness in critical applications like preference tuning and model routing.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: This paper, "Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators," is all about using activation edits to tame that internal bias within AI evaluators. The authors are Jou Barzdukas, Matthew Nguyen, Matthew Bozoukov, Simon Hongyu Fu, and Narmeen Oozeer Martian Research. They are essentially showing how we can adjust the model's "thought process" mid-judgment to make it less biased toward itself.

Jane: To put it simply, they’re proposing a way to use small modifications—these steering vectors—applied at inference time to guide the AI away from favoring its own responses when comparing different outputs. It’s a subtle intervention that doesn't require retraining the massive underlying model structure.

Lu: The implication here is that we can get closer to fairer evaluation systems for things like preference tuning and model routing because we are addressing a fundamental flaw in how these models judge each other on the fly.

Meng: But the mechanism sounds complex; what exactly are these activation edits doing inside the model to change its preference direction? We need concrete details on that process.

Lalam: I see this as a way to improve culture because if evaluators become less self-preferring, we encourage a more collaborative and unbiased interaction between different AI systems in our workflows.

The paper's summary: Tom: The core summary is that they first set up a framework to clearly separate self-preference bias from actual quality differences by using gold judges to categorize evaluations into illegitimate self-preference, legitimate self-preference, and unbiased agreement.

Jane: That separation is key because it means they aren't just trying to fix everything at once; they are specifically targeting the unfair favoritism that comes from the AI choosing its own output over another model's output unfairly.

Lu: They then introduce two main methods for creating these steering vectors: Contrastive Activation Addition, or CAA, and a data-efficient optimization method based on contrastive promotion and suppression.

Meng: So they’re not just tweaking weights randomly; they are building specific vectors by contrasting positive and negative examples of the target behavior to isolate exactly which direction in the model's activation space corresponds to that self-preference.

Lalam: That sounds like a very targeted approach, focusing on isolating the biased signal rather than trying to broadly adjust the entire model's behavior. That focus could lead to much more precise improvements in how we use these tools for safety and alignment.

The paper's improvements: Tom: The main improvement they demonstrate is that these steering vectors can successfully flip up to ninety-seven percent of those illegitimate self-preferences they identified, which is a significant correction compared to baseline methods like prompting or Direct Preference Optimization, which only showed about forty-nine percent flips.

Jane: That ninety-seven percent figure is what really stands out; it shows that this activation-based approach is much more effective at correcting the most egregious forms of bias in the evaluation process than prior techniques we’ve seen <ref:2509.03647#pg1>.

Lu: The paper suggests that these vectors can reliably shift the probability of selecting self-preference toward the mean judgment derived from impartial gold judges, and they find this effect is visible when looking at a blind versus an aware setting in their tests.

Meng: The authors also found that context-unaware vectors performed better than context-aware ones, which suggests that the representation of self-preference itself might be derivable from a relatively linear space within the model's internal structure.

Lalam: If we can indeed derive this representation from a linear space, it means we have a more structured way to understand and potentially counteract these biases in complex decision-making processes across different AI agents.

Conclusion: Tom: So, to wrap this up on "Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators," these research findings show that we can use lightweight steering vectors at inference time to reliably reduce unjustified self-preference bias by up to ninety-seven percent <ref:2509.03647#pg1>.

Jane: It really shows that we don't always need heavy training cycles like DPO just to get a strong correction; subtle activation edits can make a big difference in how AI judges things. This gives us a much more accessible tool for improving fairness in downstream applications.

Lu: The implication is that the structure of self-preference itself might be linearly encoded within the residual stream, which opens up new avenues for understanding and controlling these biases at a very low level of the model's operation.

Meng: I see this translating into practical terms as a way to build more robust model routing systems where we can trust that the evaluation pipeline isn't being subtly manipulated by internal favoritism.

Lalam: For me, this work is significant because it provides a tangible method for injecting impartiality into the very mechanism of AI judging, which could foster a much more reliable and trustworthy ecosystem for all AI interactions.

Jou Barzdukas, Matthew Nguyen, Matthew Bozoukov, Simon Hongyu Fu, Dani Roytburg

University of Virginia · University of California, San Diego · Carnegie Mellon University

cs.CL, cs.AI, cs.LG

Submitted: 2025-09-03

Updated: 2026-10-05

Comments: Presented at {Mechanistic Interpretability, Evaluations, Reliable-ML} Workshops, NeurIPS 2025

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 86/100

The gist: Large language models (LLMs) used as evaluators suffer from self-preference bias, where they disproportionately favor their own outputs, which undermines fairness in critical applications like

Key concepts

Self-Preference Bias
This occurs when an LLM evaluator disproportionately favors its own generated outputs over others. It is measured by the probability-weighted difference in selections between the model and a comparison model on a given task.
Steering Vectors
These are vectors constructed during training that modulate a model's behavior at inference time without requiring retraining. They are built using methods like Contrastive Activation Addition or optimization to steer the model toward desired behaviors.
Illegitimate Self-Preference
This refers to biased selections where the LLM favors its own output due to internal bias, rather than reflecting actual quality differences. Steering vectors were primarily effective at correcting this specific type of bias.

Terminology

Summary

Large language models (LLMs) used as evaluators suffer from self-preference bias, where they disproportionately favor their own outputs, which undermines fairness in critical applications like preference tuning and model routing. This paper investigates mitigating this issue at inference time using lightweight steering vectors to reduce unjustified self-preference bias by up to 97%.

Demonstrating Self-Preference Bias

The research first establishes a framework to disentangle self-preference bias from ground-truth quality. This is achieved by evaluating a pairwise preference set, where a model J generates summaries and a comparison model K generates other summaries. Self-preference bias is defined as the probability-weighted difference in selections averaged over the dataset: bias(J, X) = 1/X Σ i (P(vi = yJ,i) − P(vi = yK,i)). To separate this from genuine quality differences, a set of gold judges (G) from diverse model families is used to generate objective labels. This process categorizes each evaluation into three outcomes: illegitimate self-preference, legitimate self-preference, and unbiased agreement.

Constructing Steering Vectors

The authors construct steering vectors through two primary methods designed to modulate model behavior at inference time with minimal training cost. These methods are:

  1. Contrastive Activation Addition (CAA): This method builds the vector by pairing positive and negative examples for the target behavior and averaging the hidden-state activation differences they induce. The formal definition involves calculating a vector based on activations in the residual stream at a specific layer L, contrasting biased completions with unbiased ones.

  2. Optimization-based Steering: This approach learns an additive vector using a contrastive promotion/suppression method defined by [Dunefsky and Cohan, 2025] to train an additive vector with a contrastive loss function. The optimization is framed as minimizing a composite loss function that balances maximizing the probability of the desired completion against minimizing the probability of the undesired completion.

Steering Evaluations and Results

The constructed steering vectors are evaluated based on two metrics: effectiveness—the fraction of biased votes corrected—and stability—the fraction of correct votes preserved (covering unbiased agreement and legitimate self-preference). The results show that steering vectors can reliably reduce illegitimate self-preference and achieve high effectiveness, with three of the four tested vectors successfully flip up to 97% of previously biased samples. For instance, CAA achieved a flip rate of 0.97 for illegitimate self-preference. However, the paper notes that these vectors struggle with stability, as they demonstrate little modulation in legitimate self-preference and low flip rates in unbiased agreement, suggesting that self-preference spans multiple or nonlinear directions.

Model and Dataset Setup

The experiments are conducted on the XSUM dataset, a subjective summarization task consisting of 1,000 articles. The judge model used is Llama 3.1-8B-Instruct, and comparison models include GPT-3.5 OpenAI. For generating gold labels, an ensemble of models is used: Phi-4 [Abdin et al., 2024], DeepSeek V3 [DeepSeekAI et al., 2025], and Claude 3.5-Sonnet [Anthropic, 2024]. The evaluation compares the performance of steering vectors against baselines such as prompting and Direct Preference Optimization (DPO) baselines.

Discussion on Representation

The findings suggest that the disparity in stability across different preference types implies a deeper structure to self-preference. The authors conclude that illegitimate self-preference is linearly encoded while the other two are different directions either linearly or nonlinearly encoded in the residual stream. Future work is suggested to explore this possibility further and to incorporate individual rather than pairwise evaluations to circumvent issues related to ordering bias.

Applications and Context

The paper demonstrates that steering vectors can flip up to 97% of illegitimate self-preferences compared to prompting (0% flips) and DPO (49%). The results are visualized across three subsets: Bias, Agreement, and Legitimate Self-Preference. Furthermore, the authors show that context-unaware vectors outperformed their aware counterparts, suggesting that the representation of self-preference can be derived from a linear space. This work provides a safeguard against LLM-as-judges by offering an intervention at inference time with minimal training cost.


The gist

Steering vectors can reduce unjustified self-preference bias by up to 97% without retraining, though they show instability on legitimate self-preference and unbiased agreement, suggesting self-preference spans multiple or nonlinear directions in activation space.

How it works

Improvements for AI systems

Here are the specific improvements for AI systems based on the findings in this research:

  1. Improve LLM-as-Judge reliability by implementing lightweight, inference-time steering vectors (e.g., Contrastive Activation Addition or optimization-based methods) instead of relying solely on prompting or fine-tuning for bias mitigation.

  2. Implement a mechanism to distinguish between illegitimate self-preference (where the model favors its own output despite objective evidence) and legitimate self-preference (where the model correctly identifies its own superior output).

  3. Develop AI systems capable of dynamically adjusting their evaluation strategy based on the context, specifically utilizing an aware prompt setting that explicitly labels generated content as your response versus other model’s response.

  4. Create LLM evaluators that can reliably shift their probability distribution toward an impartial judge mean (the output of gold judges) by up to 97% for illegitimate biases, substantially reducing the risk of introducing model-specific preference bias into downstream tasks like preference tuning and routing.

  5. Design steering vector applications for high-stakes decision pipelines (e.g., model routing or safety checks) that ensure evaluations are robust against self-preference bias without requiring costly full retraining cycles (DPO).

Abstract

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.

Sources

Related papers