Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
cs.CL, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- Prioritizing the Best: Incentivizing Reliable Multimodal Reasoning by Rewarding Beyond Answer Correctness
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
- Understanding R1-Zero-Like Training: A Critical Perspective
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning
- HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
- Kimi K3: Open Frontier Intelligence
- TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
- GRRM: Group Relative Reward Modeling for Machine Translation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering