Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
cs.LG, cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 22 pages, 7 figures, 8 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered.
Terminology
Abstract
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau 2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V d* 0.8, whereas measured V d is 0.15 on tau 2-bench and 0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Group-in-Group Policy Optimization for LLM Agent Training
- Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Let's Verify Step by Step
- ToolACE: Winning the Points of LLM Function Calling
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration
- ToolRL: Reward is All Tool Learning Needs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- ScienceWorld: Is your Agent Smarter than a 5th Grader?
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Verifiable Process Rewards for Agentic Reasoning
- Group Sequence Policy Optimization
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks