AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing
cs.AI
Submitted: 2026-09-13
Updated: 2026-09-19
Project page: https://theappliedscientist.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce.
Terminology
Abstract
Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce. Yet a review is only useful if acting on it leads to a measurable improvement in the paper. We present AppliedScientist, a closed-loop system that couples an autonomous AI scientist with an AI reviewer, and evaluate it by iteratively revising rejected papers from a range of research subfields. To mirror how human authors build on earlier drafts, the AI scientist has access to its previous versions during revision. To avoid bias from prior judgments, however, each review is generated independently, with the reviewer having no memory of earlier feedback or scores. We compare three revision settings: one initialized with the original venue reviews, one initialized with AI-generated reviews, and autonomous self-revision using the same fixed prompt in every round. Because the reviewer both guides and evaluates the revision, we also assess the human-initialized revisions using Stanford Reviewer as an independent evaluator. Reviewer-guided revision consistently improves more than fixed-prompt self-revision, and Stanford Reviewer also assigns higher scores to later revisions. AppliedScientist resolves 128 of 150 execution-related weaknesses (85.3%), but only 2 of 18 idea-related weaknesses (11.1%), suggesting that iterative revision is effective at improving experiments and implementation, but rarely changes concerns about novelty or significance.
Sources
- ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- MARG: Multi-Agent Review Generation for Scientific Papers
- ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Kosmos: An AI Scientist for Autonomous Discovery
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions
- DeepReviewer 2.0: A Traceable Agentic System for Auditable Scientific Peer Review
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- APRES: An Agentic Paper Revision and Evaluation System
- DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process
- Endless Jailbreaks with Bijection Learning
- In-context Interference in Chat-based Large Language Models
- LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
- Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers
- One Model for One Graph: A New Perspective for Pretraining with Cross-domain Graphs
- Exploring Representations and Interventions in Time Series Foundation Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection