The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
cs.LG, cs.AI
Submitted: 2026-03-06
Updated: 2026-09-26
Comments: v3: substantially revised, new title. 30 pages, 3 figures
Code: https://github.com/EIT-EAST-Lab/C3
License: http://creativecommons.org/licenses/by/4.0/
The gist: Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the
Terminology
Abstract
Credit assignment for a team of LLM agents, what each message was worth, has mostly been treated as prediction: a learned critic, a trajectory-level score, or an agent ablation stands in for the counterfactual that defines credit, which is rarely run. Teams that communicate through a shared context are different: when everything a downstream agent reads is written into the trace, the trace is the state, and that counterfactual can be executed. A credit signal can then be judged as any estimator is: by bias, variance, and agreement with an independent reference. C3, credit assignment by counterfactual continuation, substitutes one message at a decision point and continues the run to the terminal reward, so its credit is unbiased, exact up to Monte Carlo error. Given the sampled alternatives, that error's variance follows a derived law with no term for the number of agents, and the observed noise follows the law on 6 workflows of 2 to 10 decision points. On the 2-agent chain, from 4 continuations, C3 ranks alternatives at 0.69 rank correlation against a 16-continuation reference, near the 0.73 at which 2 such references agree; a critic trained on the same rollouts reaches 0.29. Used as the advantage in training a 2-agent team, C3 beats MAGRPO, the stronger baseline in our comparison, and spends 37% fewer training tokens than MAPPO, since only the messages downstream of a decision point are regenerated. When the trace is the state, credit need not be predicted; it can be exact.
Sources
- Program Synthesis with Large Language Models
- Training Verifiers to Solve Math Word Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Qwen2.5-Coder Technical Report
- Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
- MARFT: Multi-Agent Reinforcement Fine-Tuning
- LLM Collaboration With Multi-Agent Reinforcement Learning
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Value-Decomposition Networks For Cooperative Multi-Agent Learning
- Hindsight Credit Assignment for Long-Horizon LLM Agents
- Marginal-Contribution Policy Gradients under Filtered Feedback for Multi-Agent LLMs
- QPLEX: Duplex Dueling Multi-Agent Q-Learning
- CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks