PaperGym: Rubric-Centered Evolution for Research-Plan Generation
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 34 pages, 6 figures, 6 tables. Code: https://github.com/ZJU-REAL/PaperGym. Project page: https://zju-real.github.io/PaperGym. Dataset: https://huggingface.co/datasets/CabbageWyh/PaperGym-Data. Model: https://huggingface.co/CabbageWyh/PaperGym-Model
Code: https://github.com/ZJU-REAL/PaperGym
Project page: https://zju-real.github.io/PaperGym
License: http://creativecommons.org/licenses/by/4.0/
The gist: Research planning is the decisive capability of AI scientists.
Terminology
Abstract
Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The rubric is further compressed into a single scalar per rollout. We introduce PaperGym, a unified framework that turns each research paper into a complete training environment. PaperGym exploits the structure of a paper: the question is synthesized from the research goal and background, while the criteria are derived from the method and experiments. The criteria span methodological innovation and experimental design, and criterion leakage falls to 3.7%, versus 11.90% to 34.10% in existing datasets. Training uses the rubric twice: first as privileged context for OPSD's self-teacher, then as the reward for GRPO. Across Qwen3-1.7B/4B/8B, this schedule outperforms supervised fine-tuning, either stage alone, and the reverse ordering, improving five-benchmark averages by +5.6, +5.0, and +4.8 points. With the recipe held fixed, models trained on PaperGym-20k win 58.1% of three-way comparisons, against 28.2% for RubricHub Science. The trained Qwen3-8B reaches 73.48 on ResearchQA, above the far larger Kimi K2.6. We release the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Intern-S1: A Scientific Multimodal Foundation Model
- DeepInnovator: Triggering the Innovative Capabilities of LLMs
- Training AI Co-Scientists Using Rubric Rewards
- Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
- Reinforcement Learning with Rubric Anchors
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Self-Distilled Agentic Reinforcement Learning
- SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- Reward Hacking in Rubric-Based Reinforcement Learning
- Kosmos: An AI Scientist for Autonomous Discovery
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers
- EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenAI GPT-5 System Card
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering