"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
cs.CL, cs.CR
Submitted: 2025-11-03
Updated: 2026-09-07
Comments: EMNLP 2026 Findings
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
- ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing
- Automatic and Universal Prompt Injection Attacks against Large Language Models
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Prompt Injection attack against LLM-integrated Applications
- TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
- Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
- AgentReview: Exploring Peer Review Dynamics with LLM Agents
- ReviewEval: An Evaluation Framework for AI-Generated Reviews
- PeerArg: Argumentative Peer Review with LLMs
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering