Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents

arXiv:2608.13179 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Zechuan Wang, Siyuan Lu, Hongxuan Zhang, Linjian Mo, Chenyi Zhuang, Leilei Gan

Zhejiang University · Shanghai Innovation Institute · Westlake University · AWorld Team, Inclusion AI · Nanjing University

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/openclaw/openclaw

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 95/100

The gist: The paper introduces CREST (Hierarchical Credit Assignment via Entropy-Gated Self-Teacher), a framework for training multi-turn multi-step LLM agents that addresses the hierarchical credit assignment

Terminology

Summary

The paper introduces CREST (Hierarchical Credit Assignment via Entropy-Gated Self-Teacher), a framework for training multi-turn multi-step LLM agents that addresses the hierarchical credit assignment problem in reinforcement learning with verifiable rewards (RLVR). The central research question is: Can we design a training approach that maintains the verifier-bounded performance ceiling of RL while providing token-level signal to address the credit assignment challenges of multi-turn multi-step LLM agents training?

The paper identifies a two-level credit assignment problem in multi-turn multi-step agentic RL:

  1. Inter-turn (coarse level): trajectories mix successes and failures across turns — standard RL broadcasts a single trajectory-level reward across all tokens, making credit assignment ill-posed when turns carry independent outcomes yet share one reward signal. The paper notes that on WildToolBench, no model exceeds 15% session accuracy, and performance degrades sharply with turn count.

  2. Intra-turn (fine level): individual turns conflate high-value decisions with low-entropy formatting — within a single turn, all tokens share the same advantage, treating high-value decisions identically to low-entropy formatting tokens.

  • RLVR/GRPO: standard objectives broadcast a single trajectory-level reward across all tokens — within a single turn this is merely noisy, but in multi-turn sessions it makes credit assignment ill-posed.

  • On-policy distillation (OPD): requires a same-family, tokenizer-matched teacher and its performance ceiling remains teacher-bounded.

  • On-policy self-distillation (OPSD): is brittle and restricts exploration — without per-token KL clipping, OPSD collapses within 100 training steps due to gradient concentration on a few tokens, and its performance ceiling remains teacher-bounded.

CREST uses a unified policy gradient objective with structured per-token advantages:

The objective (Equation 1): J(θ) = E[Σ t A t · log π θ(y t y<t, x)]

The per-token advantage factorization (Equation 2): A t = A turn[t] · φ t

Where:

  • A turn[t] is a turn-segmented verified-reward advantage (inter-turn credit)

  • φ t ∈ [1, 1+λϵ] is a magnitude modulation factor (intra-turn credit)

Instead of computing a single trajectory-level reward averaged across turns, we compute group-relative advantages within each turn independently. For a trajectory with turns 1,...,K, each turn k receives its own verified reward R k. The advantage for turn k in rollout i is:

A(i) k = (R(i) k − mean(R(j) k)) / (std(R(j) k) + ϵ)

This ensures that a failed turn receives negative advantage regardless of whether other turns in the same trajectory succeeded.

The modulation factor uses a privileged self-teacher (the same model conditioned on ground-truth context):

φ t = 1 + λ eff t · (w t − 1)

Key components:

  • Teacher–student divergence: Δ t = (log π T(y t h T t) − log π θ(y t h t)) / τ

  • Token weight: w t = clip(exp(sign(A turn[t]) · Δ t), 1−ϵ, 1+ϵ)

  • Direction gate: g dir t = 1[sign(A turn[t]) · Δ t > 0] — When teacher and verifier disagree on the direction of update, g dir t = 0 and the teacher signal is fully suppressed.

  • Entropy gate: uses surprisal u t = −log π θ(y t h t) with Z-score normalization and sigmoid mapping — For low-uncertainty format tokens, m ent t 1 increases it.

  • Composed gating: λ eff t = clip(λ · g dir t · m ent t, 0, λ)

The framework guarantees three formal properties:

  • (P1) Direction preservation: sign(A t) = sign(A turn[t]) always

  • (P2) Bounded bias: ∥Bias∥ ≤ λϵ · ∥∇J GRPO∥ ≈ 8.4% with λ=0.3, ϵ=0.28

  • (P3) Sign-consistent amplification: φ t ≥ 1 when active

The paper's central finding: the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes. The verified reward determines all gradient directions while the self-teacher only modulates magnitudes, ensuring the performance ceiling stays verifier-bounded.

  • BFCL V3: 100 fixed IDs from multi-turn Base split for training, 400 non-overlapping examples for evaluation (100 each from Base, Missing Functions, Missing Parameters, Long-Context)

  • WildToolBench: 256 multi-turn sessions, 128 for training, rest for evaluation

  • Qwen3-4B-Instruct (instruction-tuned)

  • Qwen3-8B (thinking model)

  • GRPO: standard trajectory-level reward

  • MT-GRPO: per-agent-step advantages

  • EnvTuning: fine-grained process rewards

  • OPD: on-policy distillation with same-family teacher

  • OPSD: on-policy self-distillation

BFCL V3 (Average accuracy):

  • Qwen3-4B: CREST achieves 52.00% vs. best baseline MT-GRPO at 49.25% (+29.88 over base)

  • Qwen3-8B: CREST achieves 50.00% vs. best baseline EnvTuning at 46.00% (+16.62 over base)

WildToolBench (Session Accuracy):

  • Qwen3-4B: CREST achieves 7.03% vs. best baseline at 6.25% (+3.90 over base)

  • Qwen3-8B: CREST achieves 9.38% vs. best baseline at 7.81% (+4.69 over base)

Key findings:

  • RL-based methods (GRPO, MT-GRPO, EnvTuning) systematically outperform distillation-based methods (OPD, OPSD) at both model scales

  • OPSD... scores 38.75%, falling below GRPO (43.63%) despite using a privileged ground-truth-conditioned teacher, giving direct empirical evidence for the teacher-bounded ceiling

  • "The margins are largest on the splits that stress hierarchical credit assignment: on BFCL Long Context... CREST improves over the strongest baseline by +7.0 on 4B and +2.0 on 8B; on WildToolBench Session Accuracy... by +0.78 on 4B and +1.57 on 8B"

  • CREST reaches 0.60 accuracy within 20 steps and converges to 0.70, while GRPO rises slowly to 0.57 by step 160

  • OPSD plateaus at 0.49 and declines thereafter

  • Gradient concentration analysis: OPSD exhibits extreme concentration: the top-1% of tokens account for 42% of the total gradient signal, and the top-5% account for 77%; GRPO is far more diffuse (top-10% ≈ 31%); CREST occupies the desirable middle ground: more concentrated than GRPO (top-10% ≈ 57%)

Hierarchical decomposition: Inter-turn segmentation alone improves average accuracy over GRPO (43.63 → 47.88)... Intra-turn modulation alone yields +5.12 points... Combining both levels achieves +8.37 points (43.63 → 52.00)

Gating ablation: Removing the direction gate drops average accuracy from 52.00% to 46.75% (−5.25)... Removing the entropy gate yields a similar degradation (52.00 → 46.25)... Removing both gates (43.50%) falls to GRPO-level performance

λ sensitivity: λ=0.3 provides the best trade-off; higher values over-concentrate gradients (λ=0.5 ends at 0.61 vs. λ=0.3 at 0.69)

The paper demonstrates that "a self-teacher need not determine gradient directions to be useful—restricting it to magnitude modulation, gated by student uncertainty, unlocks dense credit assignment while preserving the verifier-bounded ceiling that makes RL effective."

"Broader validation across additional benchmarks and larger model scales would further strengthen the generality claims. Additionally, the entropy gate uses per-token surprisal as a proxy for token importance, which is effective for separating format tokens from content tokens in tool-use trajectories but remains a heuristic that may not optimally capture token informativeness in all generation contexts."

Improvements for AI systems

Based on this paper, I can make the following specific improvements to AI systems:

1. Hierarchical Credit Assignment for Multi-Turn Agents

  • Implement turn-segmented reward advantages instead of trajectory-level rewards, so each turn in a multi-step conversation receives independent credit based on its own verified outcome

  • This enables agents to learn which specific turns caused failures, rather than attributing all errors to the entire session

2. Entropy-Gated Self-Teacher for Token-Level Signal

  • Use the model's own uncertainty (per-token surprisal) to modulate training signal strength—high-uncertainty content tokens get amplified learning signal, while low-uncertainty formatting tokens get suppressed

  • This prevents gradient waste on trivial tokens (like JSON syntax, punctuation) and focuses learning on semantically important decisions

3. Direction-Gated Privileged Teacher

  • Condition a self-teacher on ground-truth context (privileged information) but only allow it to modulate update magnitude, never direction

  • The verifier determines all gradient directions; the teacher only amplifies or dampens them when it agrees with the verifier's assessment

  • This preserves the verifier-bounded performance ceiling while providing dense, informative signal

4. Direction-Preserving Advantage Modulation

  • Guarantee that token-level advantages never contradict turn-level advantages (formal property P1)

  • This prevents conflicting learning signals that could destabilize training

A. Multi-Turn Tool-Use Agents

  • Correctly identify which specific API call or tool invocation in a multi-step workflow failed, and adjust only that behavior

  • Maintain performance on long-horizon tasks (10+ turns) without the sharp degradation seen in current systems (which drop below 15% session accuracy on complex benchmarks)

B. Reasoning and Code Generation Models

  • Distinguish between high-value reasoning steps and low-value formatting tokens during RL training

  • Learn faster (reaching target accuracy in 20 steps vs 160 steps for standard RL) while maintaining final performance

C. Systems with Limited Verification Signals

  • When only sparse, coarse rewards are available (e.g., final answer correctness), still provide dense per-token learning signal through the entropy-gated self-teacher

  • This works even without per-step reward models or process supervision

D. Distillation-Compatible Systems

  • Enable on-policy self-distillation without the collapse problem (which occurs within 100 steps in current methods)

  • Achieve this through per-token KL clipping and direction gating, allowing stable self-improvement without external teachers

E. Systems Requiring Controlled Exploration

  • Maintain exploration diversity through the bounded bias guarantee (≤8.4% deviation from pure RL)

  • Avoid the gradient concentration problem that causes OPSD collapse (top-1% tokens accounting for 42% of gradient signal) while still focusing learning on important tokens (top-10% ≈ 57% vs GRPO's diffuse 31%)

F. General-Purpose Agent Training Pipelines

  • Combine the best of both worlds: RL's verifier-bounded ceiling with distillation's dense token-level signal

  • Achieve consistent improvements across model scales (4B and 8B) and task types (function calling, tool use, long-context reasoning) without requiring task-specific reward engineering

Abstract

Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce CrEST, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. CrEST resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that CrEST consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.

Sources

Related papers