What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal
cs.SE, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/radinshayanfar/task_snc
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- On Randomness in Agentic Evals
- FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
- Evaluating Large Language Models Trained on Code
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
- Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- Agentic Software Issue Resolution with Large Language Models: A Survey
- The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering
- SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
- SWE-smith: Scaling Data for Software Engineering Agents
- SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
- When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties