Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics
cs.CL
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: Accepted at CIKM 2026
Code: https://github.com/ReviewerlyInc/llm_judge_reliability_analysis
License: http://creativecommons.org/licenses/by/4.0/
The gist: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk.
Terminology
Abstract
Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.
Sources
- PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
- ReviewEval: An Evaluation Framework for AI-Generated Reviews
- ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing
- PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
- Verbosity Bias in Preference Labeling by Large Language Models
- OpenAI GPT-5 System Card
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering