SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
cs.SE, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: 11 pages, 2 figures, 5 tables
Code: https://github.com/DeepSoftwareAnalytics/SWE-Gate
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- CoCoSum: Contextual Code Summarization with Multi-Relational Graph Neural Network
- What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
- PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
- RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
- HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
- Knowledge Matters: Injecting Project and Testing Knowledge into LLM-based Unit Test Generation
- Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties