Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
cs.MA, cs.AI, cs.SE
Submitted: 2026-09-07
Updated: 2026-09-28
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
- Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
- Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements
- Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
- Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning