Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
cs.CL, cs.AI, cs.IR, cs.LG, cs.MA, cs.SE
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/zjunlp/AutoSciRub
Terminology
Sources
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape
- Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation
- LightMem-Ego: Your AI Memory for Everyday Life
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- Qworld: Question-Specific Evaluation Criteria for LLMs
- ResearchGym: Evaluating Language Model Agents on Real-World AI Research
- Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
- Autonomous LLM-driven research from data to human-verifiable research papers
- RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
- Kosmos: An AI Scientist for Autonomous Discovery
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering