Marginal-Contribution Policy Gradients under Filtered Feedback for Multi-Agent LLMs
cs.LG
Submitted: 2026-04-03
Updated: 2026-09-08
Comments: 22 pages, 1 figure. Substantially revised version
License: http://creativecommons.org/licenses/by/4.0/
The gist: We develop a unified treatment of credit assignment for RL training in multi-agent LLM systems.
Terminology
Abstract
We develop a unified treatment of credit assignment for RL training in multi-agent LLM systems. We show that observed reward alone cannot distinguish an agent that determines it from one that never affects it, and that standard shared-reward training performs exact gradient ascent on each agent's private utility rather than system performance. Moreover, we prove no single scalar per agent can consistently account for joint performance once agents interact. We thus develop the unique background-dependent notion of marginal contribution satisfying natural consistency requirements. From it we derive gradient-correct marginal contribution training signals, identify them from filtered feedback, and optimally allocate a budget of exact counterfactual evaluations against learned-signal error. Instantiated in GRPO, our signal improves routed GSM8K accuracy over winner-take-all training at no extra generation cost.
Sources
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- CollabLLM: From Passive Responders to Active Collaborators
- Doubly Robust Policy Evaluation and Learning
- Training Verifiers to Solve Math Word Problems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks