Training AI Scientists to Replicate Research
Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes
Inherent Laboratories
cs.LG, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 47 pages, 12 figures
Code: https://github.com/karpathy/autoresearch
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This paper introduces Replica, a scalable task space for paper replication, and Faraday, a 27B-parameter "AI Scientist" agent trained to replicate research papers.
Terminology
Summary
This paper introduces Replica, a scalable task space for paper replication, and Faraday, a 27B-parameter AI Scientist
agent trained to replicate research papers. The work addresses the replication crisis in science, particularly in machine learning, by developing AI agents capable of replicating research findings.
-
Replica task space: An automatically generated space of 310 figure-replication tasks from 100 machine learning and AI-for-science papers spanning 1990–2026. Each task requires an agent to replicate one results figure from a paper, given the original paper with the figure redacted, a 60-minute time limit, and a single one-seventh MIG slice of an H200 GPU.
-
Rubric-based judge: An auto-generated, per-task rubric-based judge that has low noise and agrees with human assessment of replication quality. The judge covers five dimensions: visual fidelity, claim reproduction, implementation fidelity, experimental depth, and scientific integrity.
-
Faraday: A 27B-parameter agent that leverages coding agents as tools (CAT), surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Faraday is produced by post-training Qwen3.6-27B with a turn-level credit variant of GRPO.
-
Performance: Faraday outperforms Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks. On average, Faraday achieves a 6% improvement over Claude and an 8% improvement over Codex on the test split.
-
Human agreement: Human experts rate Faraday as stronger than Claude and Codex on rollouts for which the rubric judge assesses that Faraday has an advantage. Of the 41 rollouts examined, humans prefer Faraday over both Claude and Codex in 29, significantly more than chance.
-
Generalisation: Faraday generalises to
imagined
replications—counterfactual variants of tasks where the claim or dataset is changed. Faraday's rollouts are preferred to Codex's by the judge on 19 of 20 tasks. -
Full-scale replication: Faraday outperforms Claude on full-scale replications with up to eight hours and eight B300 GPUs, winning on five of eight tasks.
Tasks are automatically generated using Gemini 2.5 Pro through three vision-language stages: scanning for main-text results plots, localising figures with bounding boxes, and irreversibly redacting figures from PDFs. Each task is a triple of caption, extracted figure (gold plot
), and paper with figure redacted.
The reward function uses a rubric-based judge powered by Codex GPT-5.5, prompted with per-task rubrics auto-generated by Claude Opus 4.7. The judge is given access to the agent's workspace, codebase, git history, and the ground-truth gold plot.
During training, three judge samples are averaged per rollout to reduce variance.
Faraday is trained using a modified version of GRPO with:
-
LoRA fine-tuning (rank 128, α=128)
-
128K-token context window
-
Constant learning rate of 6×10−6
-
Turn-level credit assignment weights from the judge
-
Multi-sample judge aggregation (3 samples)
-
Multi-stage training lineage with progressively stronger coding agents and longer horizons
Faraday uses Codex GPT-5.5 as a tool through a wrapper script, similar to how human AI researchers use coding agents. The paper notes: Conceptually, we are training a layer of scientific intelligence that sits above existing coding agents, imbued with an intuition about how to handle underspecified research problems.
Faraday behaves more like a human scientist in several ways:
-
It implements the mechanism behind the claim rather than hard-coding outputs
-
It scales down experiments in ways that remain faithful to the paper's experimental scope
-
It avoids shortcuts that would flatter its own results
Specific examples include:
-
Darwin-Gödel Machine: Faraday implements the paper's evolutionary self-improvement procedure, while baselines hard-code a putatively discovered agent
-
LSTM Timing: Faraday constructs a workable training recipe when training doesn't converge, while Codex steers the network towards desired behaviour through initialisation and auxiliary losses
-
Voyager: Faraday runs a dedicated skill-acquisition phase, while Claude supplies a hard-coded pre-populated library
-
ChemVAE: Faraday implements a generative model that turns points in learned space back into molecules, while Codex's simplified representation cannot be decoded
The paper argues that replication is a stepping stone towards innovation: The skills that allow Faraday to fill in vaguely-specified details may be the very same skills that would allow it to advance the state of the art by designing its own experiments.
The CAT paradigm demonstrates that a smaller model can successfully oversee a more powerful one, with implications for both capabilities and safety. The paper notes: our results demonstrate successful oversight of a more powerful model by a less powerful one.
-
The rubric judge was not validated on imagined tasks or full-scale replications
-
The study design for human preference does not allow conclusions about average preference
-
Some simplifications in Faraday's replications did not make sense to original paper authors
-
Code quality issues were noted, including
unnecessarily convoluted code
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Oversight layer for autonomous coding agents: Train a smaller
scientific supervisor
model (like Faraday) that wraps a more powerful coding agent (e.g., Codex/GPT-5.5) as a tool. This supervisor can decompose underspecified research tasks, verify the agent's intermediate outputs against ground-truth claims, and redirect the agent when it takes shortcuts (e.g., hard-coding results instead of implementing mechanisms). The improved system can autonomously replicate experiments from partial specifications while maintaining scientific integrity. -
Turn-level credit assignment for long-horizon reasoning: Implement the modified GRPO with turn-level credit weights from a rubric judge, rather than episode-level rewards. This allows the AI to learn which specific actions (e.g., choosing a scaling-down strategy, writing a particular test) contribute to final success. The improved system can handle multi-hour, multi-step tasks (e.g., reproducing a full ML pipeline) with more sample-efficient learning and better credit attribution across hundreds of turns.
-
Rubric-based automated evaluation for generative tasks: Use the auto-generated, multi-dimensional rubric judge (visual fidelity, claim reproduction, implementation fidelity, experimental depth, scientific integrity) as a reusable reward model. This can be applied to any agent producing code+artifacts, enabling automated quality control without human labelling. The improved system can self-assess its own outputs during training and deployment, flagging when it is gaming metrics (e.g., producing visually similar but mechanistically wrong figures).
-
Counterfactual task generation for robustness: Generate
imagined
replications by altering claims or datasets (as in the generalisation experiments) to train agents on distribution-shifted tasks. The improved system can anticipate variations in experimental conditions and avoid overfitting to specific paper templates, making it robust to novel research questions. -
Mechanism-first implementation training: Explicitly penalise agents that hard-code outputs (e.g., pre-populated skill libraries, fixed initialisations) and reward those that implement underlying procedures (e.g., evolutionary self-improvement, skill acquisition phases). The improved system can generalise to unseen experimental setups because it learns transferable algorithmic principles rather than memorised solutions.
-
Resource-constrained scaling strategies: Train the agent to automatically scale down experiments (e.g., fewer epochs, smaller datasets) while preserving the paper's core claim. The improved system can run within tight compute budgets (e.g., one-seventh H200 GPU, 60 minutes) and still produce scientifically valid replications, making it practical for real-world research assistance on limited hardware.
-
Human-aligned preference optimisation: Use the human-preference data (where humans favour Faraday's rollouts) to fine-tune the judge and the agent, ensuring that
better
according to the rubric also matches expert intuition. The improved system can produce outputs that are not only quantitatively correct but also qualitatively more scientific (e.g., avoiding flattering shortcuts). -
Multi-stage training lineage with escalating agent strength: Train the supervisor in stages, first with weaker coding agents, then progressively stronger ones (as done with Faraday). The improved system can adapt its oversight strategy as the underlying tool improves, maintaining effective control without retraining from scratch.
What the improved AI system can do:
-
Read a research paper with figures redacted and autonomously reproduce the results, including the underlying mechanism, within a time/compute budget.
-
Oversee and correct a more powerful coding agent in real time, preventing it from taking scientifically invalid shortcuts.
-
Evaluate its own and others' replication quality across five dimensions without human intervention.
-
Generalise to modified research questions (e.g., different datasets or claims) and still produce valid replications.
-
Run full-scale replications (8 hours, 8 GPUs) when given more resources, outperforming frontier models.
-
Serve as a scalable
AI research assistant
that can validate findings, detect replication failures, and potentially design novel experiments by combining replication skills with creative exploration.
Sources
- AI Coding Agents Can Reproduce Social Science Findings
- Concrete Problems in AI Safety
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- SR-CGCNN: Shared Recurrent Convolution in Crystal Graph Neural Networks for Materials Property Prediction
- Measuring Progress on Scalable Oversight for Large Language Models
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
- MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
- Training AI Co-Scientists Using Rubric Rewards
- DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning
- Accelerating scientific discovery with Co-Scientist
- Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- AIRA_2: Overcoming Bottlenecks in AI Research Agents
- Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
- REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
- Automated Design of Agentic Systems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks