Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
cs.SE, cs.AI, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- SWE-Bench+: Enhanced Coding Benchmark for LLMs
- An Information-Theoretic Perspective on Credit Assignment in Reinforcement Learning
- Intrinsic Credit Assignment for Long Horizon Interaction
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
- Surprisal-Guided Selection: Compute-Optimal Test-Time Strategies for Execution-Grounded Code Generation
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
- Process Reinforcement through Implicit Rewards
- TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
- SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
- MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
- Communication-Aware Multi-Agent Reinforcement Learning for Decentralized Cooperative UAV Deployment
- SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
- Scaling Test-Time Compute for Agentic Coding
- Learning Game-Playing Agents with Generative Code Optimization
- Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation
- Understanding the Challenges in Iterative Generative Optimization with LLMs
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties