LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
cs.SE, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/SWE-agent/mini-swe-agent
Terminology
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval
- Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
- ReAct: Synergizing Reasoning and Acting in Language Models
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
- FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties