Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
cs.AI, stat.ML
Submitted: 2026-09-17
Updated: 2026-09-17
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a
Terminology
Abstract
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.
Sources
- Process Reinforcement through Implicit Rewards
- Mind2Web: Towards a Generalist Agent for the Web
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- AgentBench: Evaluating LLMs as Agents
- Understanding R1-Zero-Like Training: A Critical Perspective
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Self-Refine: Iterative Refinement with Self-Feedback
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- ScienceWorld: Is your Agent Smarter than a 5th Grader?
- EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection