Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
cs.CL
Submitted: 2026-02-02
Updated: 2026-09-09
Comments: Accepted to Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge.
Terminology
Abstract
Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examined point-wise or pair-wise evaluation protocols; in contrast, our focus is on rubric-based evaluation, which has been attracting increasing attention owing to its utility for training models in domains where verification is otherwise difficult. In this work, we show that rubric-based evaluation implicitly resembles a multiple-choice setting and therefore exhibits position bias: LLMs tend to prefer score options that appear at specific positions within the rubric list. Through controlled experiments across multiple models and datasets, we demonstrate that this position bias is consistent. Its direction, however, is model-specific: some judges favor the first option, while others favor the last. We further identify a second, orthogonal axis of bias: when a prompt scores several criteria simultaneously, the ordering of the criteria itself shifts the resulting scores. We additionally explore permuting the order of the rubric options as a means of mitigating position bias, and find that although the bias can be attenuated, improvements in the correlation between model judgments and human annotations are obtained primarily for models that exhibit strong bias. Our results recast rubric-based LLM-as-a-judge as a multiple-choice problem with measurable, model-specific position bias, and we further confirm that only a small number of random order permutations are sufficient to reduce the error introduced by this bias for the majority of models.
Sources
- News Summarization and Evaluation in the Era of GPT-3
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Do LLM Evaluators Prefer Themselves for a Reason?
- Evaluating Scoring Bias in LLM-as-a-Judge
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- gpt-oss-120b & gpt-oss-20b Model Card
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
- Gemma 3 Technical Report
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering