Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
cs.SE, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/earino/identical-runs-different-results
Project page: https://szilard.github.io/xgboost-autoresearch
Terminology
Sources
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- On Randomness in Agentic Evals
- Accounting for Variance in Machine Learning Benchmarks
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
- Evaluating Large Language Models Trained on Code
- Generalization in Adaptive Data Analysis and Holdout Reuse
- An Empirical Study of Harness Design for Coding Agents
- Can LLMs Beat Classical Hyperparameter Optimization Algorithms? A Study on autoresearch
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Towards a Science of AI Agent Reliability
- When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties