ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
cs.CL, cs.AI, cs.LG
Submitted: 2026-06-16
Updated: 2026-09-01
Comments: EMNLP 2026 Main Conference
Code: https://github.com/LithiumDA/ReproRepo
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reproducing research results from papers and released code is central to scientific progress.
Terminology
Abstract
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility, but they are difficult to scale due to their reliance on substantial manual effort for data curation and evaluation. We introduce ReproRepo, a scalable framework for reproducibility evaluation that leverages human-raised GitHub issues as naturally occurring supervision on realistic reproduction blockers. We instantiate ReproRepo on 1,149 recent machine learning papers from major conferences and evaluate four frontier model-agent configurations. Our results show that LLM agents, even without executing code, can identify many real-world reproducibility problems from paper-repository pairs: the best agent in our study, namely Codex with GPT-5.5, surfaces at least one semantically related human-reported blocker for about 90% of papers in the study. Further analysis shows that agents are particularly effective for surfacing visible failures and identifying the right semantic region, but may still be insufficient in exact localization. ReproRepo can serve as a reusable, scalable framework for future evaluations of LLM agents on real-world reproducibility auditing. Our code is released at https://github.com/LithiumDA/ReproRepo.
Sources
- ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
- Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
- Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Replication and Reanalysis
- Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results
- ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
- ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
- SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
- FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
- Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
- PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering