LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation
Han Li, Zhemin Fang, Rili Feng, Yingqi Zhao, Jiaheng Liu, Pengfei Gao, He Ye, Dayi Lin, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
cs.SE, cs.CL
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Project page: https://loopsbench.ai/
Project page: https://loopsbench.ai
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
- FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
- Evaluating Large Language Models Trained on Code
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- RExBench: Can coding agents autonomously implement AI research extensions?
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
- Measuring Coding Challenge Competence With APPS
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Measuring AI Ability to Complete Long Software Tasks
- Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
- Training Software Engineering Agents and Verifiers with SWE-Gym
- Qwen2.5 Technical Report
- CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
- OmniCode: A Benchmark for Evaluating Software Engineering Agents
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties