Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time

arXiv:2601.15172 · cs.CL · Submitted 2026-01-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time".

Jane: The gist The cross-temporal analysis reveals no consistent decline in median review quality across venues and years.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: The paper introduces a new framework for this evidence-based comparison of review quality, applying it to ICLR, NeurIPS, and ACL conferences. It’s about moving beyond just looking at one conference in isolation.

Jane: They are proposing a way to measure review quality as utility for both the editors and the authors themselves using these specific criteria—substantiveness, actionability, and grounding.

Lu: The authors point out that reviewing practices evolve constantly, which creates this diversity in how reviews look across different venues and time periods.

Meng: They address this structural mess by developing a novel approach to unify these reviews into a common format using Large Language Models to preserve the meaning while fixing the structure.

Lalam: So they flatten the diverse review structures into continuous text, which is then itemized into bullet point-style items so you can measure things consistently.

The paper's summary: Tom: The core of what they found in this study is that when you look across time and different venues, there isn't a consistent decline in the median review quality.

Jane: That’s counter-intuitive to what some people assume, so the paper does a detailed cross-temporal analysis to check that assumption.

Lu: They used both lightweight measurements like review length and request count alongside more complex LLM-based metrics like actionability and grounding scores for their analysis.

Meng: They found that while there are fluctuations in quality, there isn't a steady, substantial drop across the years they looked at.

Lalam: The overall finding is that the median review quality at computer science conferences has not declined in recent years, even when you look at the data from ICLR, NeurIPS, and ACL.

The paper's improvements: Tom: The authors suggest a few ways to take this framework further. They are proposing five specific hypotheses that researchers can test to dig into these nuanced changes over time.

Jane: They recommend specific things for researchers and conference organizers on how they can use this new multi-dimensional schema to monitor quality better in real time.

Lu: One improvement they suggest is integrating this framework into a continuous monitoring pipeline so you can flag low-quality reviews as they happen at any venue or time.

Meng: They also suggest using LLM-based unification to create that common form, which should help ensure the measurements aren't skewed by structural inconsistencies between different conference guidelines.

Lalam: And they propose developing lightweight approximations for those heavier LLM metrics so that you can actually scale the analysis and screen a lot of reviews quickly.

Conclusion: Tom: So, to wrap up, the main implication is that we shouldn't be worried about a consistent quality drop in peer review based on this study of "Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time."

Jane: They’ve shown that while things fluctuate, there isn't a steady decline when you use this holistic framework to look at the data from ICLR, NeurIPS, and ACL.

Lu: The paper sets up a path for researchers to formally test if these quality metrics are actually improving or changing as the field grows.

Meng: From an engineering standpoint, having a unified way to measure utility across different review styles is really useful for building better tools in this area.

Lalam: I think it's helpful because it gives everyone a standardized way to talk about what makes a review good, using those dimensions of substantiveness and actionability.

Ubiquitous Knowledge Processing Lab (UKP Lab) · Department of Computer Science, Technical University of Darmstadt · National Research Center for Applied Cybersecurity ATHENE · Department of Computer Science at Queens College, City University of New York (CUNY)

cs.CL

Submitted: 2026-01-21

Updated: 2026-10-08

Comments: To appear in EMNLP-2026

Code: https://github.com/UKPLab/arxiv2026-review-quality-estimation

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: The gist The cross-temporal analysis reveals no consistent decline in median review quality across venues and years.

Key concepts

Review Quality Schema
This is a multi-dimensional system for judging a review based on three main aspects: substantiveness (how much useful information is conveyed), actionability (how easy it is to act upon the feedback), and grounding (how well the review relates to the original work). These dimensions are measured using specific metrics.
Review Unification Approach
This method solves the problem of reviews having different structures. It uses a Large Language Model (LLM) to take varied review formats and convert them into a single, continuous text form, allowing for consistent comparison across different conferences and time periods.
Lightweight vs. LLM-based Measurements
The study uses two types of metrics: simple, easy-to-deploy measurements like review length (LEN) and request count (REQ), alongside more complex metrics generated by LLMs, such as the Actionability score (ACT) and Grounding score (GND). This combination allows for both broad and deep analysis of review quality trends.
Cross-Temporal Analysis
This involves examining how review quality changes across different years and various conferences. The analysis showed that while there are fluctuations, there is no consistent or substantial decline in the median review quality observed over time.

Terminology

Summary

The gist The cross-temporal analysis reveals no consistent decline in median review quality across venues and years.

Framework for Evidence-Based Comparative Study

The paper introduces a new framework for evidence-based comparative study of review quality and applies it to major AI and machine learning conferences: ICLR, NeurIPS and ACL<ref:2601.15172A, We propose a multi-dimensional schema for quantifying review quality as utility to editors and authors, coupled with both LLM-based and lightweight measurements.

Review Quality Schema

The proposed schema is multi-dimensional, spanning three high-level criteria: substantiveness, actionability, and grounding<ref:2601.15172B, We develop a novel review quality schema that spans multiple dimensions – substantiveness, actionability, and grounding. These criteria are operationalized through specific measurements like ITX for Substantiveness and ACT for Actionability<ref:2601.15172C, Substantiveness captures the overall amount of information conveyed in a review and serves both authors and editors.

Review Unification Approach

A key challenge addressed is the diversity of review structures across conferences and time, which is solved by a novel LLM-based approach for meaning-preserving review structure unification<ref:2601.15172D, We identify challenges in unifying reviewing data across conferences and time, and develop a novel LLM-based approach for meaningpreserving review structure unification. This involves flattening reviews into a continuous text (r f) and then itemizing them into self-contained bullet point-style items (r i)<ref:2601.15172E, To unify reviews produced using differently structured forms, we flatten the reviews by concatenating all the text fields and section headers from a review form into a continuous text (Figure 3, top), r → r f.

Measurements and Analysis

The study utilizes both lightweight measurements—such as Review length (LEN) and Request count (REQ)—and LLM-based metrics like Actionability score (ACT) and Grounding score (GND)<ref:2601.15172F, We propose a set of lightweight measurements that are easy to deploy, and we explore LLM-based measurements that enable deeper analysis. The framework allows for the study of relationships between these measurements and their evolution over time<ref:2601.15172G, We apply this schema to study the interactions between different measurements and to analyse the changes in review quality at computer science conferences over time.

Conclusion on Trends

The cross-temporal analysis reveals no consistent decline in median review quality across venues and years<ref:2601.15172H, Contrary to the widespread perception, we find that the median review quality at computer science conferences shows no consistent decline. The authors propose five hypotheses to further investigate this nuanced question<ref:2601.15172I, We provide five hypotheses coupled with recommendations to researchers, conference organizers and community leaders to move this work forward. The study concludes that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years.

The paper introduces a new framework for evidence-based comparative study of review quality and applies it to major AI and machine learning conferences: ICLR, NeurIPS and ACL<ref:2601.15172A, We propose a multi-dimensional schema for quantifying review quality as utility to editors and authors, coupled with both LLM-based and lightweight measurements. The framework is built around a multi-dimensional schema spanning substantiveness, actionability, and grounding<ref:2601.15172B, We develop a novel review quality schema that spans multiple dimensions – substantiveness, actionability, and grounding. Review unification is achieved through flattening reviews into a continuous text (r f) and then itemizing them into self-contained bullet point-style items (r i)<ref:2601.15172E, To unify reviews produced using differently structured forms, we flatten the reviews by concatenating all the text fields and section headers from a review form into a continuous text (Figure 3, top), r → r f. The study applies this schema to study interactions between different measurements and changes in review quality over time<ref:2601.15172G, We apply this schema to study the interactions between different measurements and to analyse the changes in review quality at computer science conferences over time. The analysis uses lightweight metrics like Review length (LEN) and Request count (REQ) alongside LLM-based metrics like Actionability score (ACT) and Grounding score (GND)<ref:2601.15172F, We propose a set of lightweight measurements that are easy to deploy, and we explore LLM-based measurements that enable deeper analysis. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The paper concludes by proposing five hypotheses to further investigate this nuanced question<ref:2601.15172I, We provide five hypotheses coupled with recommendations to researchers, conference organizers and community leaders to move this work forward. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The final finding is that the median review quality at computer science conferences has not declined in recent years<ref:2601.15172J, Although we observe some fluctuations in review quality, there is no consistent and substantial decline across years. The study's conclusion is that the median review quality at computer science conferences has not declined in recent: 9. Conclusion"

How it works

  1. A new framework for evidence-based comparative study of review quality is introduced, which includes a multi-dimensional schema for quantifying review quality as utility to editors and authors<ref:2601.15172B, We propose a multi-dimensional schema for quantifying review quality as utility to editors and authors, coupled with both LLM-based and lightweight measurements.

  2. A novel approach is developed for meaning-preserving review structure unification using LLMs to transform reviews into a common form<ref:2601.

Improvements for AI systems

  1. The proposed framework for a holistic evaluation of review quality can be integrated into a continuous monitoring pipeline to detect low-quality reviews in real-time across different venues and time periods. This system will utilize the multi-dimensional schema (substantiveness, actionability, and grounding) to flag reviews that consistently score poorly on utility metrics like ITX or ACT.

  2. Implement an LLM-based review structure unification module to transform diverse review formats into a common form, mitigating the challenges identified in Section 4. This will allow for consistent application of quality measurements across conferences with varying guidelines, ensuring that the LLM-based measurements are not skewed by structural inconsistencies.

  3. Develop lightweight, efficient measurement approximations for LLM-based metrics to enable scalable analysis. Since lightweight measurements can approximate LLM-based measurements (Section 7.2), this will allow for rapid screening of large volumes of reviews without incurring the high computational cost associated with full LLM processing.

  4. Create a dynamic review quality score based on a non-weighted average of all our measurements. This aggregate score can be used as a leading indicator for tracking the overall health of the peer-review ecosystem, providing editors and authors with an objective metric to assess the current state of review quality over time.

  5. Establish hypothesis testing protocols using bootstrap tests on median Q values across years and venues. This will allow researchers to formally test hypotheses like H1 (Review quality didn’t decline) with statistical rigor, moving beyond anecdotal evidence to establish causal links between interventions and review quality improvements.

Abstract

Peer review is at the heart of modern science. As submission numbers rise and research communities grow, the decline in review quality is a popular narrative and a common concern. Yet, is it true? Review quality is difficult to measure, and the ongoing evolution of reviewing practices makes it hard to compare reviews across venues and time. To address this, we introduce a new framework for evidence-based comparative study of review quality and apply it to major AI and machine learning conferences: ICLR, NeurIPS and *ACL. We document the diversity of review formats and introduce a new approach to review standardization. We propose a multi-dimensional schema for quantifying review quality as utility to editors and authors, coupled with both LLM-based and lightweight measurements. We study the relationships between measurements of review quality, and its evolution over time. Contradicting the popular narrative, our cross-temporal analysis reveals no consistent decline in median review quality across venues and years. We propose alternative explanations, and outline recommendations to facilitate future empirical studies of review quality.

Sources

Related papers