A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/c-steve-wang/Robust_Review
Terminology
Sources
- Stop Automating Peer Review Without Rigorous Evaluation
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent
- GLM-5: from Vibe Coding to Agentic Engineering
- Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
- Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
- How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- gpt-oss-120b & gpt-oss-20b Model Card
- Large Language Models are Inconsistent and Biased Evaluators
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- CycleResearcher: Improving Automated Research via Automated Review
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering