Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
Yifu Huo, Shunjie Xing, Chenglong Wang, Peinan Feng, Qiaozhi He, Yan Ding, Anxiang Ma, Yuxin Gao, Tongran Liu, Tong Xiao, Jingbo Zhu
cs.LG, cs.CL
Submitted: 2026-08-08
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments.
Terminology
Abstract
Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.
Sources
- Reinforcement Learning for Long-Horizon Interactive LLM Agents
- Training Verifiers to Solve Math Word Problems
- Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
- Group-in-Group Policy Optimization for LLM Agent Training
- Multimodal Web Navigation with Instruction-Finetuned Foundation Models
- A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
- SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
- Sparse Rewards Can Self-Train Dialogue Agents
- Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
- Let's Verify Step by Step
- A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
- Qwen2.5 Technical Report
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Hindsight Credit Assignment for Long-Horizon LLM Agents
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks