SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
cs.SE, cs.LG
Submitted: 2026-01-20
Updated: 2026-09-08
Code: https://github.com/Aider-AI/aider
License: http://creativecommons.org/licenses/by/4.0/
The gist: Software testing is crucial for ensuring the correctness and reliability of software systems.
Terminology
Abstract
Software testing is crucial for ensuring the correctness and reliability of software systems. Automated generation of issue reproduction tests from natural language issue descriptions enhances developer productivity by simplifying root cause analysis, promotes test-driven development -- "test first, write code later", and can be used for improving the effectiveness of automated issue resolution systems like coding agents. Existing methods proposed for this task predominantly rely on closed-source LLMs, with limited exploration of open models. To address this, we propose SWE-Tester -- a novel pipeline for training open-source LLMs to generate issue reproduction tests. First, we curate a high-quality training dataset of 41K instances from 2.6K open-source GitHub repositories and use it to train LLMs of varying sizes and families. The fine-tuned models achieve absolute improvements of up to 10% in success rate and 21% in change coverage on SWT-Bench Verified. Further analysis shows consistent improvements with increased inference-time compute, more data, and larger models. These results highlight the effectiveness of our framework for advancing open-source LLMs in this domain.
Sources
- TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
- Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- LocAgent: Graph-Guided LLM Agents for Code Localization
- LoRA: Low-Rank Adaptation of Large Language Models
- Qwen2.5-Coder Technical Report
- AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
- Decoupled Weight Decay Regularization
- Issue2Test: Generating Reproducing Test Cases from Issue Reports
- Training Software Engineering Agents and Verifiers with SWE-Gym
- Gemma 3 Technical Report
- AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
- SWE-smith: Scaling Data for Software Engineering Agents
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- AutoCodeRover: Autonomous Program Improvement
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties