How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
University of Maryland · Virginia Tech · MBZUAI · University of Waterloo
cs.CL, cs.AI
Submitted: 2026-08-10
Updated: 2026-08-11
Code: https://github.com/MingLiiii/Dissecting_AI_Reviews
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper investigates how rhetorical choices in scientific manuscripts can influence AI-based peer review judgments, a form of reward hacking where presentation changes scores without altering the
Terminology
Summary
The paper investigates how rhetorical choices in scientific manuscripts can influence AI-based peer review judgments, a form of reward hacking where presentation changes scores without altering the underlying science. The authors construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters (GPT-5.5 and Opus 4.8) transform six rhetorical dimensions in opposing directions, and five LLM reviewers (Gemini 3.5 FL, Qwen 3.5 F, GPT-5 mini, GPT-5.5, Sonnet 5) evaluate the resulting manuscripts under standard and strict protocols. The study also tests joint, recursive, and reviewer-guided rewriting workflows.
The results show that rhetorical sensitivity is structured rather than uniform. The paper's key findings are:
Finding 1: Rewrite sensitivity concentrates in specific dimensions and varies with the AI reviewer’s initial score. "Evidence framing and novelty stance produce the largest effects, followed by scope framing: positive evidence framing raises Overall Assessments (OA) by up to 0.93, negative novelty stance lowers it by up to 0.73, and evidence framing changes weak-accept probability by 13 percentage points on average. This hierarchy is stable across human-assessed quality levels, but
lower initial AI scores tend to rise, higher scores tend to fall, and directional contrasts are strongest in the middle ranges."
Finding 2: More elaborate rewriting yields configuration-dependent and diminishing gains. "Joint rewriting produces substantial gains with Opus 4.8 but near-zero gains with GPT-5.5. Reviewer guidance does not consistently outperform an unguided second pass at the same depth, and gains diminish after the second pass."
Finding 3: Rewriters shape rhetorical contrasts, reviewers shape rewrite effects, and protocols shift the score scale. "Rewriter mainly changes the separation between positive and negative versions, whereas the reviewer changes both the magnitude and the sign of the resulting OA effects. Strict review lowers absolute OA by 1.36 points on average but does not consistently strengthen or weaken rewrite effects. The reviewer also determines whether OA changes are reflected primarily in contribution or soundness."
The paper concludes that "AI scientific review is systematically sensitive to rhetorical presentation even when reported scientific content is preserved. This sensitivity is not uniform. It depends on the rhetorical dimension, the rewriting process, the rewriter and reviewer models, and the review protocol. The resulting variation cannot be reduced to a single average effect or treated as a stable property of one model configuration. The findings suggest that
AI-assisted review should be evaluated for rhetorical robustness across multiple models and conditions, rather than judged solely by aggregate agreement or average scoring behavior."
Improvements for AI systems
Improvements to AI systems:
-
Add rhetorical-robustness auditing to AI review systems. The improved system can automatically detect and flag when a manuscript’s score shifts by more than a threshold (e.g., 0.5 OA) due to changes in evidence framing, novelty stance, or scope framing, even when the scientific content is identical. It can generate a “rhetorical sensitivity report” per dimension, per reviewer model, and per protocol, enabling users to see if their review is being gamed.
-
Implement score-calibration layers that correct for presentation bias. The improved system can apply a learned correction factor based on the rewriter model and reviewer model pair, normalizing OA scores to reduce the 0.93-point inflation from positive evidence framing and the 0.73-point deflation from negative novelty stance. It can output a “content-only score” that strips rhetorical variance, using the paper’s finding that sensitivity is concentrated in specific dimensions.
-
Create multi-model ensemble review with rhetorical variance weighting. The improved system can run the same manuscript through multiple reviewer LLMs (e.g., Gemini 3.5 FL, Qwen 3.5 F, GPT-5 mini) and weight their scores inversely to their measured rhetorical sensitivity, so that a reviewer who is less swayed by presentation gets higher weight. It can also flag when reviewers disagree in sign or magnitude of rewrite effects, indicating unstable judgments.
-
Add a “rhetorical stress test” mode for authors and editors. The improved system can automatically generate positive and negative versions of a manuscript’s evidence framing and novelty stance (using the paper’s controlled rewriting dimensions), then show how the AI review score changes. This lets authors see if their paper’s acceptance probability is fragile, and lets editors reject submissions that only pass due to presentation.
-
Build a reviewer-guidance system that reduces reward hacking. The improved system can, during the review process, explicitly instruct the reviewer LLM to ignore rhetorical cues (e.g., “do not adjust score based on how evidence is framed”) and then compare the guided score to the unguided score. If the difference exceeds a threshold, the system can flag the review as potentially hacked and request a re-review under a strict protocol that lowers absolute OA by 1.36 points but does not amplify rewrite effects.
-
Develop a recursive rewriting monitor for submission integrity. The improved system can detect when a manuscript has undergone joint or reviewer-guided rewriting (by analyzing stylistic fingerprints from the paper’s finding that gains diminish after the second pass). It can then automatically re-evaluate the manuscript under a strict protocol and report whether the score change is due to content improvement or rhetorical manipulation, preventing inflated acceptances.
-
Create a dimension-specific scoring breakdown for transparency. The improved system can decompose each OA score into contributions from evidence framing, novelty stance, scope framing, and other dimensions, based on the paper’s finding that the reviewer determines whether changes appear in contribution or soundness. This allows users to see exactly which rhetorical aspect drove the score, and to compare across reviewers to identify systematic biases.
-
Implement adaptive review protocols based on initial score. The improved system can use the paper’s finding that lower initial AI scores tend to rise and higher scores tend to fall under rewriting. It can automatically apply a stricter protocol for manuscripts with mid-range scores (where directional contrasts are strongest) to reduce variance, and a more lenient protocol for extreme scores, improving decision reliability.
Sources
- Stop Automating Peer Review Without Rigorous Evaluation
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
- Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
- Large Language Models are Inconsistent and Biased Evaluators
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering