Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 10 pages, full research paper, to appear in proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM '26), 2026
Code: https://github.com/infobip/llm-judge-
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm.
Terminology
Abstract
Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgment is often assumed to be independent of earlier evaluations. We test this assumption using three prompt conditions: no metadata, revision framing, and anchored metadata containing revision, attempt, and prior-score fields. We show that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values. Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts. Cohen's d, a standardized measure of the difference between score distributions, reaches an absolute value of 0.71. Token-level analysis of selected model-task probes suggests a threshold-like response pattern: introducing anchored metadata produces a marked redistribution of output-score probabilities, while changing the anchor value within the tested below-threshold range produces comparatively little additional variation. On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to categorical decisions. Neither Chain-of-Thought nor a metadata-disregard warning reduces the total effect, although the warning improves the paired accuracy effect relative to baseline in the industry experiment. Reliable LLM evaluation demands careful context engineering rather than an assumption of impartiality. Effective mitigation must be validated for the intended model and task or domain.
Sources
- Qwen2.5 Technical Report
- STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
- Gemma 2: Improving Open Language Models at a Practical Size
- Personalized Prediction of Perceived Message Effectiveness Using Large Language Model Based Digital Twins
- Understanding the Anchoring Effect of LLM with Synthetic Data: Existence, Mechanism, and Potential Mitigations
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- The Llama 3 Herd of Models
- Large Language Models are Inconsistent and Biased Evaluators
- Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption
- Self-Preference Bias in LLM-as-a-Judge
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- Wider and Deeper LLM Networks are Fairer LLM Evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering