When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
cs.CL, cs.AI
Submitted: 2025-09-30
Updated: 2026-08-30
License: http://creativecommons.org/licenses/by/4.0/
The gist: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM as a benchmark) where a model generates test inputs (LLM as a testset) and evaluates outputs (LLM as an
Terminology
Abstract
As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM as a benchmark) where a model generates test inputs (LLM as a testset) and evaluates outputs (LLM as an evaluator) has gained traction as a cheap alternative to human curation. We show that this paradigm has a fundamental problem: LLM-generated benchmarks systematically favor the model that created them. Using machine translation as our primary testbed, we find that self bias arises from two compounding sources, LLM as a testset and LLM as an evaluator, and their combination amplifies the effect. Crucially, even when test data is generated with explicit diversity controls, each modelś implicit stylistic tendencies produce homogeneous, model-specific outputs that inflate its own scores. Increasing source text diversity, using our proposed diversity metric, partially mitigates this bias. Self bias is strong enough to cause each model to rank itself first, overriding the peer consensus ordering. We confirm that the phenomenon extends to open-ended generation on the Chatbot Arena task.
Sources
- Do LLM Evaluators Prefer Themselves for a Reason?
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Efficacy of Synthetic Data as a Benchmark
- Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
- FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
- Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Large Language Models Are State-of-the-Art Evaluators of Translation Quality
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
- Automated Benchmark Generation for Repository-Level Coding Tasks
- Self-Preference Bias in LLM-as-a-Judge
- Large Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias
- Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-Generator
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering