RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
cs.AI, cs.LG, cs.SE
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/mithils3/reclaim
Terminology
Sources
- Improving Generalization of Neural Combinatorial Optimization for Vehicle Routing Problems via Test-Time Projection Learning
- MirrorCode: AI can rebuild entire programs from behavior alone
- From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- AI Coding Agents Can Reproduce Social Science Findings
- SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
- Recovering Wasted Compute in Autoresearch Agents
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP
- RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
- Exploring the use of AI authors and reviewers at Agents4Science
- On Randomness in Agentic Evals
- SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
- Accounting for Variance in Machine Learning Benchmarks
- You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
- Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
- Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection