Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
cs.CL, cs.AI
Submitted: 2026-06-06
Updated: 2026-09-06
Comments: EMNLP Findings 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability.
Terminology
Abstract
As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narrative constraints selected from a shared pool, which carry no stylistic fingerprint yet retain a recoverable model-specific signature. Two experiments on this task yield complementary findings. Under blind evaluation, self-preference disappears, with a small effect remaining in the opposite direction once selection quality and judge severity are controlled. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally. LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
Sources
- Do LLM Evaluators Prefer Themselves for a Reason?
- Style over Story: Measuring LLM Narrative Preferences via Structured Selection
- Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
- Self-Preference Bias in LLM-as-a-Judge
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
- Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
- Self-critiquing models for assisting human evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering