Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
cs.CL, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted for publication at EMNLP 2026. 5 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric.
Terminology
Abstract
LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.
Sources
- The Llama 3 Herd of Models
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Reference-free Evaluation Metrics for Text Generation: A Survey
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Qwen2.5 Technical Report
- Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
- Gemma: Open Models Based on Gemini Research and Technology
- MedGemma 1.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering