Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models

arXiv:2609.33812 · cs.SE, cs.LG · Submitted 2026-09-27 · Read on arXiv

cs.SE, cs.LG

Submitted: 2026-09-27

Updated: 2026-09-27

Code: https://github.com/earino/identical-runs-different-results

Project page: https://szilard.github.io/xgboost-autoresearch

Terminology

Sources

Related papers