Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
Chaoran Chen, Vy Nguyen, Ziji Zhang, Abhinav Gullapalli, Ziyi Wang, Yuxuan Lu, Dakuo Wang, Jing Huang, Zhou Yu, Jin Lai
Amazon
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper presents BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly
Terminology
Summary
The paper presents BENCH2ROBUST, a framework that converts failure-free tool-use benchmarks into controlled stochastic environments with scenario-controlled solvability, where episodes explicitly require retrying, switching, or stopping after available paths are exhausted. The framework is used to study two complementary interventions: structured runtime recovery context through Bayesian Tool Memory (BTM) and curriculum-controlled reinforcement learning. Across 7 models from 4 families and two multi-turn benchmark families, tool failures produce a near-universal robustness gap. On held-out Retail tasks, BTM improves robustness by up to 16.8 percentage points without retraining, while RL learns complementary recovery behavior that remains beneficial without inference-time BTM. Combining the two reaches 40.8–45.5% under injection while preserving failure-free performance.
The paper addresses the problem that Most tool-use benchmarks, however, assume that tool calls return timely, structurally valid, correct, and current responses.
Production tools violate these assumptions: requests time out or hit rate limits, credentials and schemas become invalid, and apparently valid responses may contain stale, partial, or incorrect data.
Robust agents must do more than persist after failure. Depending on what remains recoverable, they may need to RETRY the same path, SWITCH to an alternative source, or ABSTAIN and escalate after viable paths are exhausted.
A central challenge is that conventional stochastic failure injection does not identify which recovery behavior an episode actually requires.
The paper introduces scenario-controlled solvability with three recoverability structures:
-
S1 (retry works):
no path is permanently blocked and retrying can recover
-
S2 (switch needed):
the primary path is blocked episode-wide and completion requires a fallback-equivalent tool
-
S3 (impossible):
all available paths are blocked, defining the environment state in which escalation is appropriate
BENCH2ROBUST is a benchmark-agnostic injection framework that interposes between agent and tool execution.
It formalizes the injected environment as a POMDP with a corrupted observation kernel. Each tool call independently samples a noise type with P(clean) = 0.60; the remaining 0.40 is spread over 9 failure modes organized by observability: 6 explicit-signal failures (timeout, rate limit, server error, auth error, malformed, schema drift) and 3 silent corruptions (partial data, stale values, factual errors).
Scenario blocking is deterministic and episode-persistent and operates outside the stochastic injection budget (B ≤ 5, Kmax = 2 consecutive stochastic failures on the same tool).
For Retail, 5 alternative tools are added (21 total), each providing a different query pattern to the same data (e.g., search product by name as alternative to get product details).
BTM combines three components:
-
Beta-posterior recoverability statistics estimated from training-split rollouts
-
domain-specific fallback maps with granularity annotations
-
heuristic constraints such as retrying before abandoning a path, verifying information before irreversible actions, and escalating only after available recovery paths have been considered
The paper clarifies: The Bayesian component is the estimation and accumulation of empirical recoverability statistics, not the fallback structure itself.
For each (tool j, error type e) pair, recovery probability is modeled as prec j,e ∼ Beta(αj,e, βj,e) with posterior mean P̂ = α/(α + β).
The reward function is:
R(τ) = 1.0 · eff(τ) · rep(τ) if task completed
0.3 · matched actions/total required · rep(τ) otherwise
where eff(τ) = max(0.3, 1.0 − 0.02 · max(0, τ − 12)) penalizes long episodes and rep(τ) ∈ 0, 0.5, 1.0 penalizes repetition (≥4 identical calls → 0).
Importantly, the reward contains no positive term for correct abstention: S3 episodes cannot pass the benchmark task evaluator.
Five phases progressively introduce harder scenarios:
-
Phase 1: 100% S1 (RETRY), max 2 injections
-
Phases 2–3: introduce SWITCH (20%→40% S2)
-
Phase 4: adds ABSTAIN (10% S3)
-
Phase 5: full repertoire (30/45/25)
DAPO is used with a KL constraint (coefficient 0.02) to prevent policy collapse—without KL, training collapses within 25–30 iterations as repetition rate exceeds 80%.
Additional components include "asymmetric clipping (ϵlow = 0.20, ϵhigh = 0.28), dynamic group filtering (σ < 0.01 masked), and group-relative advantages over G = 16 rollouts per prompt (8 clean + 8 injected, unified normalization)."
Training and evaluation uses Combined Retail: 1,339 multi-turn customer-service tasks composed of τ2-bench retail (114 tasks) and Retail-3I (1,225 tasks covering general/ambiguous/changing intents), split into 937 training and 402 held-out tasks.
The primary training target is Qwen3-4B-Thinking-2507,
with an additional Qwen3-8B
replication. Reference models include Qwen3-8B, Qwen3-32B, Qwen3-235B, DeepSeek-V3, GLM-4.7, and MiniMax-M2.5, spanning 4 model families and scales from 4B to 235B parameters.
Of 70 (model, subset) pairs, 69 degrade (−1.8 to −46.7pp).
The gap appears across customer service, airline booking, and function-calling domains. "Larger models sometimes lose fewer absolute points, but scaling does not eliminate brittleness: Qwen3-235B still loses 15.8pp on τ2-retail, while models with 89–91% clean Telecom performance (GLM-4.7, MiniMax-M2.5) lose 20–21pp."
Key results from Table 1 (402 held-out tasks, 5 seeds):
-
Base model: Clean 64.3±0.9%, Inj w/o Alt 20.1±1.4%, Inj w/ Alt 31.9±2.3%
-
Base + BTM: Clean 65.3±0.9%, Inj w/o Alt 36.9±1.2%, Inj w/ Alt 43.6±1.6%
-
RL (no BTM at inference): Clean 64.5±0.8%, Inj w/o Alt 26.4±1.2%, Inj w/ Alt 38.8±2.2%
-
RL + BTM: Clean 63.9±1.3%, Inj w/o Alt 40.8±1.6%, Inj w/ Alt 45.5±3.0%
Structured recovery context provides the largest zero-training gain. BTM improves robustness without retraining: 20.1%→36.9% (+16.8pp w/o alt) and 31.9%→43.6% (+11.7pp w/ alt) on held-out tasks.
RL retains recovery gains without inference-time BTM. The RL-trained policy reaches 26.4%/38.8% when BTM is removed at inference, compared with 20.1%/31.9% for the base model (+6.3/+6.9pp).
Full system preserves failure-free performance. RL+BTM clean performance (63.9±1.3%) is within 0.4pp of the base model (64.3±0.9%).
From Table 2, on S2 (switch-required) episodes: RL's realized S2 success increases from 16.8% to 35.3% relative to Base+BTM, while premature escalation decreases from 52.5% to 41.7%.
The paper notes: "Purely increasing retry persistence would not be expected to help once a path is permanently blocked, so the simultaneous S2 improvement and reduction in premature escalation are consistent with recovery behavior beyond retry frequency."
From Table 3: The gain is overwhelmingly structural: 'Static only' (constraints + fallback maps with uniform P = 0.50) recovers +10.5pp of the full +14.6pp; calibrated beliefs without structure add only +1.7pp.
Under fixed beliefs, Shuffled values (38.3%) nominally exceed True (37.8%), and under online updating the margin is ≤1.6pp—at or below seed variance.
From Table 4:
-
BTM helps most on explicit transient errors:
For timeout, rate limit, server error, and malformed responses, adding BTM improves pass rate by +18.0 to +25.3pp.
-
RL contributes larger marginal gains on persistent observable errors:
On the two persistent observable errors—auth error and schema drift—the observable-group asymmetry reverses: BTM's increment drops from +18–25pp to +12.4/+12.6pp, while RL's marginal increment rises to +10.2/+12.5pp.
-
RL adds further gains on silent errors:
For factual error (+10.5pp), stale (+5.4pp), and partial (+4.5pp), there is no error signal at all.
BTMhurts on factual error (−4.6pp), the only negative entry in the table.
From Appendix C: "The Retail-trained RL model transfers to unseen domains without retraining. RL alone (without target-domain BTM) already contributes +2.7pp on Airline (33.4%) and +1.5pp on BFCL (15.0%); combining it with target-domain BTM reaches +9.0pp and +5.0pp respectively."
Qwen3-8B replication (Appendix B): "The 8B results reproduce the qualitative decomposition observed on the 4B target. BTM alone improves injected performance from 33.3% to 39.3% without alternatives and from 39.4% to 41.2% with alternatives. RL without inference-time BTM improves the same conditions to 37.4% and 42.2%, respectively. Combining RL with inference-time BTM yields the highest pass rates, 41.4% and 44.6%."
Training method comparison (Appendix M): Our complete scenario-structured recipe substantially outperforms a vanilla-GRPO random-noise baseline (45.9% vs. 23.0%; Table 16).
Sensitivity to injection severity (Appendix I): Across the clean-probability sweep, both systems improve as injection becomes milder, but RL+BTM retains a consistent advantage over the base model (+9.0 to +22.2pp).
The paper concludes: "We presented BENCH2ROBUST, a framework for studying tool-failure robustness as a recovery-policy problem rather than persistence alone. Scenario-controlled solvability separates retry-, switch-, and abstain-required regimes, exposing a broad robustness gap across seven models. On held-out Retail tasks, structured recovery context gives the largest immediate gain (+11.7 to +16.8pp), RL retains additional robustness without inference-time context (+6.3/+6.9pp), and their combination reaches 40.8–45.5% under injection without measurable clean-task degradation."
The paper suggests a division of labor: runtime context supplies environment knowledge, while training makes retry/switch behavior more reusable across episodes and, to a smaller extent, across domains.
Key limitations acknowledged:
-
our intervention comparison evaluates complete recipes rather than isolated components
-
the reward signal for correct abstention remains incomplete
-
the failures are simulated stressors rather than real incident traces
-
recovery under injection costs more tokens (70K→117K), which may matter in latency-sensitive deployments
-
Telecom is near the 4B model's capability floor (21% clean), and Retail-to-BFCL transfer is modest (+5pp)
-
S3 episodes define when escalation is appropriate, but Eq. 2 assigns no completion bonus to correct abstention
Improvements for AI systems
Based on the paper, here are specific improvements to AI systems and what the improved systems can do:
Improvement: Add a Bayesian Tool Memory (BTM) module that maintains per-tool, per-error-type recovery probability estimates (Beta posteriors) and domain-specific fallback maps, injected as structured context at inference time.
Improved capability: The system can distinguish between transient errors (retry-appropriate) and persistent errors (switch-appropriate), improving tool-use robustness by +16.8pp without any retraining, while preserving clean-task performance.
Improvement: Train policies using a 5-phase curriculum that progressively introduces retry-only → switch-required → abstain-required scenarios, with DAPO and asymmetric clipping (εlow=0.20, εhigh=0.28), KL constraint (0.02), and group-relative advantages over 16 rollouts.
Improvement: Combine BTM at inference with RL-trained recovery policy, where BTM provides environment knowledge and RL provides reusable retry/switch behavior.
Improvement: Implement a two-tier recovery strategy: BTM for explicit transient errors (timeout, rate limit, server error, malformed — +18–25pp) and RL for persistent observable errors (auth error, schema drift — +10–12pp) and silent corruptions (factual error, stale, partial — +4.5–10.5pp).
Improvement: Classify each episode into S1 (retry-works), S2 (switch-needed), or S3 (impossible) based on tool-blocking patterns, and condition the policy's action selection on this classification.
Improvement: Incorporate repetition penalties (≥4 identical calls → reward 0) and episode-length penalties into the reward function during training.
Improvement: Train on combined multi-domain data (Retail + Airline + BFCL) with domain-specific fallback maps, using the scenario curriculum across all domains.
Improvement: Provide a configuration interface that lets operators set injection severity (clean probability), maximum consecutive failures (Kmax), and alternative tool availability, then automatically selects the optimal recovery strategy (BTM-only, RL-only, or combined).
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
- DeepSeek-V3 Technical Report
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions
- When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
- Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
- Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
- Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
- ToolCritic: Detecting and Correcting Tool-Use Errors in Dialogue Systems
- Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection