Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents
cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/yaojh18/RLER
Terminology
Sources
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Process Reward Models for LLM Agents: Practical Framework and Directions
- When Agents go Astray: Course-Correcting SWE Agents with PRMs
- Tmax: A simple recipe for terminal agents
- PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
- Simulating Environments with Reasoning Models for Agent Training
- ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- The Verification Horizon: No Silver Bullet for Coding Agent Rewards
- A Rubric-Supervised Critic from Sparse Real-World Outcomes
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
- AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
- From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering