Training AI Scientists to Replicate Research

arXiv:2608.13331 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes

Inherent Laboratories

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 47 pages, 12 figures

Code: https://github.com/karpathy/autoresearch

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper introduces Replica, a scalable task space for paper replication, and Faraday, a 27B-parameter "AI Scientist" agent trained to replicate research papers.

Terminology

Summary

This paper introduces Replica, a scalable task space for paper replication, and Faraday, a 27B-parameter AI Scientist agent trained to replicate research papers. The work addresses the replication crisis in science, particularly in machine learning, by developing AI agents capable of replicating research findings.

  1. Replica task space: An automatically generated space of 310 figure-replication tasks from 100 machine learning and AI-for-science papers spanning 1990–2026. Each task requires an agent to replicate one results figure from a paper, given the original paper with the figure redacted, a 60-minute time limit, and a single one-seventh MIG slice of an H200 GPU.

  2. Rubric-based judge: An auto-generated, per-task rubric-based judge that has low noise and agrees with human assessment of replication quality. The judge covers five dimensions: visual fidelity, claim reproduction, implementation fidelity, experimental depth, and scientific integrity.

  3. Faraday: A 27B-parameter agent that leverages coding agents as tools (CAT), surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Faraday is produced by post-training Qwen3.6-27B with a turn-level credit variant of GRPO.

  • Performance: Faraday outperforms Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks. On average, Faraday achieves a 6% improvement over Claude and an 8% improvement over Codex on the test split.

  • Human agreement: Human experts rate Faraday as stronger than Claude and Codex on rollouts for which the rubric judge assesses that Faraday has an advantage. Of the 41 rollouts examined, humans prefer Faraday over both Claude and Codex in 29, significantly more than chance.

  • Generalisation: Faraday generalises to imagined replications—counterfactual variants of tasks where the claim or dataset is changed. Faraday's rollouts are preferred to Codex's by the judge on 19 of 20 tasks.

  • Full-scale replication: Faraday outperforms Claude on full-scale replications with up to eight hours and eight B300 GPUs, winning on five of eight tasks.

Tasks are automatically generated using Gemini 2.5 Pro through three vision-language stages: scanning for main-text results plots, localising figures with bounding boxes, and irreversibly redacting figures from PDFs. Each task is a triple of caption, extracted figure (gold plot), and paper with figure redacted.

The reward function uses a rubric-based judge powered by Codex GPT-5.5, prompted with per-task rubrics auto-generated by Claude Opus 4.7. The judge is given access to the agent's workspace, codebase, git history, and the ground-truth gold plot. During training, three judge samples are averaged per rollout to reduce variance.

Faraday is trained using a modified version of GRPO with:

  • LoRA fine-tuning (rank 128, α=128)

  • 128K-token context window

  • Constant learning rate of 6×10−6

  • Turn-level credit assignment weights from the judge

  • Multi-sample judge aggregation (3 samples)

  • Multi-stage training lineage with progressively stronger coding agents and longer horizons

Faraday uses Codex GPT-5.5 as a tool through a wrapper script, similar to how human AI researchers use coding agents. The paper notes: Conceptually, we are training a layer of scientific intelligence that sits above existing coding agents, imbued with an intuition about how to handle underspecified research problems.

Faraday behaves more like a human scientist in several ways:

  • It implements the mechanism behind the claim rather than hard-coding outputs

  • It scales down experiments in ways that remain faithful to the paper's experimental scope

  • It avoids shortcuts that would flatter its own results

Specific examples include:

  • Darwin-Gödel Machine: Faraday implements the paper's evolutionary self-improvement procedure, while baselines hard-code a putatively discovered agent

  • LSTM Timing: Faraday constructs a workable training recipe when training doesn't converge, while Codex steers the network towards desired behaviour through initialisation and auxiliary losses

  • Voyager: Faraday runs a dedicated skill-acquisition phase, while Claude supplies a hard-coded pre-populated library

  • ChemVAE: Faraday implements a generative model that turns points in learned space back into molecules, while Codex's simplified representation cannot be decoded

The paper argues that replication is a stepping stone towards innovation: The skills that allow Faraday to fill in vaguely-specified details may be the very same skills that would allow it to advance the state of the art by designing its own experiments.

The CAT paradigm demonstrates that a smaller model can successfully oversee a more powerful one, with implications for both capabilities and safety. The paper notes: our results demonstrate successful oversight of a more powerful model by a less powerful one.

  • The rubric judge was not validated on imagined tasks or full-scale replications

  • The study design for human preference does not allow conclusions about average preference

  • Some simplifications in Faraday's replications did not make sense to original paper authors

  • Code quality issues were noted, including unnecessarily convoluted code

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Oversight layer for autonomous coding agents: Train a smaller scientific supervisor model (like Faraday) that wraps a more powerful coding agent (e.g., Codex/GPT-5.5) as a tool. This supervisor can decompose underspecified research tasks, verify the agent's intermediate outputs against ground-truth claims, and redirect the agent when it takes shortcuts (e.g., hard-coding results instead of implementing mechanisms). The improved system can autonomously replicate experiments from partial specifications while maintaining scientific integrity.

  2. Turn-level credit assignment for long-horizon reasoning: Implement the modified GRPO with turn-level credit weights from a rubric judge, rather than episode-level rewards. This allows the AI to learn which specific actions (e.g., choosing a scaling-down strategy, writing a particular test) contribute to final success. The improved system can handle multi-hour, multi-step tasks (e.g., reproducing a full ML pipeline) with more sample-efficient learning and better credit attribution across hundreds of turns.

  3. Rubric-based automated evaluation for generative tasks: Use the auto-generated, multi-dimensional rubric judge (visual fidelity, claim reproduction, implementation fidelity, experimental depth, scientific integrity) as a reusable reward model. This can be applied to any agent producing code+artifacts, enabling automated quality control without human labelling. The improved system can self-assess its own outputs during training and deployment, flagging when it is gaming metrics (e.g., producing visually similar but mechanistically wrong figures).

  4. Counterfactual task generation for robustness: Generate imagined replications by altering claims or datasets (as in the generalisation experiments) to train agents on distribution-shifted tasks. The improved system can anticipate variations in experimental conditions and avoid overfitting to specific paper templates, making it robust to novel research questions.

  5. Mechanism-first implementation training: Explicitly penalise agents that hard-code outputs (e.g., pre-populated skill libraries, fixed initialisations) and reward those that implement underlying procedures (e.g., evolutionary self-improvement, skill acquisition phases). The improved system can generalise to unseen experimental setups because it learns transferable algorithmic principles rather than memorised solutions.

  6. Resource-constrained scaling strategies: Train the agent to automatically scale down experiments (e.g., fewer epochs, smaller datasets) while preserving the paper's core claim. The improved system can run within tight compute budgets (e.g., one-seventh H200 GPU, 60 minutes) and still produce scientifically valid replications, making it practical for real-world research assistance on limited hardware.

  7. Human-aligned preference optimisation: Use the human-preference data (where humans favour Faraday's rollouts) to fine-tune the judge and the agent, ensuring that better according to the rubric also matches expert intuition. The improved system can produce outputs that are not only quantitatively correct but also qualitatively more scientific (e.g., avoiding flattering shortcuts).

  8. Multi-stage training lineage with escalating agent strength: Train the supervisor in stages, first with weaker coding agents, then progressively stronger ones (as done with Faraday). The improved system can adapt its oversight strategy as the underlying tool improves, maintaining effective control without retraining from scratch.

What the improved AI system can do:

  • Read a research paper with figures redacted and autonomously reproduce the results, including the underlying mechanism, within a time/compute budget.

  • Oversee and correct a more powerful coding agent in real time, preventing it from taking scientifically invalid shortcuts.

  • Evaluate its own and others' replication quality across five dimensions without human intervention.

  • Generalise to modified research questions (e.g., different datasets or claims) and still produce valid replications.

  • Run full-scale replications (8 hours, 8 GPUs) when given more resources, outperforming frontier models.

  • Serve as a scalable AI research assistant that can validate findings, detect replication failures, and potentially design novel experiments by combining replication skills with creative exploration.

Sources

Related papers