Divergent strategies and convergent outcomes in autonomous materials discovery
cs.AI, cond-mat.mtrl-sci
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/jihankim929/replicate-study
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns.
Terminology
Abstract
Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized sessions of one model-harness configuration received a frozen database of 12,499 metal-organic frameworks, a methane-storage objective, a pinned protocol and a one-week budget. Strategies diverged into four approaches spanning 100--5,000 screened structures, and eight built 2,253 hypothetical structures. Yet the agents recovered the same materials frontier near 200 cm 3/cm cubed, and an independent calculation of the database's porous region found its nine best structures all among their reports. Enforced checks on half the agents raised fresh-run reproduction from one of eight to eight of eight but could not detectably improve conclusion validity, because fifteen of sixteen agents selected the same audit-excluded entry, an incomplete structure whose missing anions created artificial pore volume. Replicated agents thus reveal both robust conclusions and common-mode errors from shared inputs.
Sources
- SimMOF: AI agent for Automated MOF Simulations
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Accelerating scientific discovery with Co-Scientist
- El Agente Quntur: A research collaborator agent for quantum chemistry
- CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
- ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
- The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
- Can Coding Agents Reproduce Findings in Computational Materials Science?
- PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection