Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
cs.SE, cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/DeepSoftwareAnalytics/PTA-IRT
Terminology
Sources
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties