Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

arXiv:2608.11552 · cs.CL, cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Dylan Bouchard, Mohit Singh Chauhan

cs.CL, cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: The paper "Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents" investigates whether uncertainty quantification (UQ) methods developed for single-turn language

Terminology

Summary

The paper Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents investigates whether uncertainty quantification (UQ) methods developed for single-turn language model outputs transfer to the interactive, multi-turn trajectory setting of LLM agents. The authors note that "Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome."

The study evaluates three common families of single-turn UQ methods across five LLMs (Qwen2.5-7B, gpt-oss-20b, Qwen3.5-9B, MiniMax-M3, and gpt-4o-mini) and four multi-turn tool-use datasets (the multi-turn subset of BFCL-v4 and the retail, airline, and telecom text datasets from τ2-bench). The three families are: white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory.

The core finding is that transfer is often useful but uneven. Specifically, the authors find that "Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. Importantly, no family is uniformly reliable across all models and domains in our evaluations."

The paper makes three main contributions. First, it provides a controlled empirical comparison of these UQ families in multi-turn tool-use agents, reporting discrimination, calibration, and selective prediction. Second, it adapts single-turn methods to trajectory-level scoring through cross-turn aggregation of token-probability scores and new measures of black-box consistency, including a model-based trajectory-equivalence scorer. Third, it finds that transferability is mixed across scoring methods.

For white-box scorers, the paper evaluates sequence probability (SP), length-normalized sequence probability (LNSP), and average token negentropy (ATN@K), applying them over the action span of each turn and aggregating across turns using various rules (first, mean, min, last, early-weighted, late-weighted). The results show that White-box performance varies with both base score and aggregator. For example, on BFCL-v4, mean or minimum entropy reaching 0.65–0.71 AUROC for gpt-4o-mini, Qwen2.5-7B, and MiniMax-M3, but on the airline dataset, the mean aggregator applied to SP and LNSP... is approximately at or below chance for every model. The paper highlights the instability: changing the aggregation rule produces large AUROC swings within model-dataset columns, and no aggregation is predictive everywhere. Across the white-box grid, 91 of 240 cells fall below 0.5, and 21 have bootstrap intervals lying entirely below 0.5.

For black-box consistency scorers, the paper evaluates final-message consistency (NCP), action-structured consistency (FAC, ASC, ADC, AEC), and judged consistency (TER). The results show that NCP ranges from 0.832 for Qwen3.5-9B on airline to at or below chance on telecom for several models. The action-structured scorers perform well on BFCL-v4, reaching about 0.71–0.77 AUROC for four models. Across the variants, TER and ASC most often occupy the top ranks. However, no single consistency scorer dominates across datasets.

For reflexive scorers, the paper evaluates P(True) and Verbalized Confidence (VC). The results show that Reflexive scoring has its strongest AUROC point estimates on BFCL-v4, ranging from 0.746–0.852 across models for P(True). Across the τ2-bench datasets, the best reflexive scores are 0.710 on retail, 0.852 on telecom, and 0.703 on airline. The paper notes that Airline shows the widest cross-model spread, with P(True) ranging from 0.225 to 0.659 and VC from 0.400 to 0.703.

The paper also includes robustness checks. A native tool-calling ablation on retail with gpt-4o-mini shows that the results do not appear to be an artifact of the text-action interface, with nearly identical TER and AEC AUROC. A prefix-conditioned ablation on airline, which holds the transcript fixed and resamples only the next agent action, finds only modest ranking ability (up to 0.588 AUROC), suggesting that downstream trajectory dynamics contributing additional signal beyond local next-action stability.

In terms of cost, the paper summarizes that "White-box scores add no extra model calls when action-token log probabilities are available, reflexive scoring adds one self-evaluation pass, and consistency scorers incur the cost of the m resampled trajectories, with TER's gemini-flash-lite judge adding a negligible 0.002 per task. The main cost is the need to run additional agent trajectories."

The paper concludes that trajectory-adapted variants of single-turn UQ have mixed performance and that agent UQ is not just single-turn UQ with a longer context. The authors recommend per-deployment validation rather than transfer by default and suggest future work on step-level labels that enable evaluation by action type and tool family, whether acting on per-turn scores improves outcomes during execution, and exploring different types of agentic environments exposed by other benchmarks (e.g., web-browsing).

Improvements for AI systems

Improvements to AI Systems:

  1. Add trajectory-level uncertainty scoring to agent frameworks. Instead of attaching confidence only to final answers, compute uncertainty per trajectory by aggregating token-probability scores across action turns (e.g., mean, min, early-weighted) and by resampling full trajectories for consistency checks. This enables the agent to flag low-confidence trajectories for human review or fallback strategies.

  2. Implement a hybrid uncertainty selector that adapts per deployment. Since no single UQ family dominates across models and domains, build a meta-scorer that, at deployment time, evaluates a small validation set and picks the best-performing scorer (e.g., among token-probability aggregators, self-consistency variants like trajectory-equivalence (TER) or action-set consistency (ASC), and reflexive P(True)). This prevents reliance on a default method that may be at chance level (e.g., mean aggregator on airline data).

  3. Use trajectory-equivalence (TER) as a default black-box consistency scorer. TER, which uses a judge model to compare resampled trajectories for equivalence, consistently ranks among the top consistency variants across datasets. Integrate it as a plug-in scorer for agent systems where resampling is affordable, with the judge cost being negligible (0.002 per task).

  4. Add per-turn action-token entropy as a real-time signal for early stopping or tool-selection adjustment. For white-box models with access to log probabilities, compute average token negentropy over the action span of each turn. Use a threshold on this per-turn score to decide whether to ask a clarifying question, re-plan, or escalate to a human—rather than waiting for final outcome uncertainty.

  5. Implement a reflexive self-assessment pass for low-cost uncertainty. Use P(True) or verbalized confidence after trajectory completion as a cheap baseline (one extra model call). In systems where resampling is too expensive, default to P(True) but validate its AUROC on a small domain-specific set, since its performance varies widely (e.g., 0.225–0.659 on airline).

  6. Build a selective prediction mechanism for agent outputs. Use the best-performing trajectory-level scorer to abstain from acting when uncertainty exceeds a threshold. For example, on BFCL-v4, reflexive P(True) reaches 0.746–0.852 AUROC, enabling reliable abstention on low-confidence trajectories, which improves overall task success rate by avoiding cascading errors from intermediate decisions.

  7. Add a prefix-conditioned uncertainty check for local action stability. Before committing to a next action, resample only the next action given the fixed transcript (not full trajectory). While this yields modest AUROC (up to 0.588), it can be combined with full-trajectory scores to catch early errors cheaply, especially in tool-use domains where downstream dynamics add signal.

  8. Incorporate domain-specific calibration. Since aggregator choice causes large AUROC swings (e.g., 91 of 240 white-box cells below 0.5), the improved system should automatically calibrate uncertainty scores per domain using a small labeled set, rather than assuming transferability. This ensures the agent does not silently degrade in new environments (e.g., telecom vs. retail).

  9. Enable cost-aware uncertainty budgeting. The system should dynamically choose between white-box (free), reflexive (1 extra call), and consistency (m resampled trajectories) scorers based on the task’s risk level and available compute. For high-stakes tasks, use TER with m=5–10; for low-stakes, use P(True) or token-probability aggregators.

  10. Add trajectory-level uncertainty to agent logs for post-hoc analysis. Store per-turn entropy, consistency scores, and reflexive confidence for each completed trajectory. This enables offline auditing, detection of systematic failure modes (e.g., specific tool families), and continuous improvement of the UQ selector without retraining the agent.

Sources

Related papers