One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents
cs.SE, cs.CL, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 44 pages, including appendices. Model available at https://huggingface.co/Logics-MLLM/Logics-SWE-Qwen3.6-27B
Code: https://github.com/kimjune01/swebench-pro-audit
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Agentic Reinforced Policy Optimization
- Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
- MiniLLM: On-Policy Distillation of Large Language Models
- Reinforcement Learning via Self-Distillation
- Editing Models with Task Arithmetic
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SWE-Prot'eg'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
- Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
- Auditing Reward Hackability in Code RL Training Environments
- Self-Improving Large Language Models via Progressive Experience Evolution
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties