When Should Multi-Round RAG Stop? Structured Stopping Judgments and Retrieval Reduction in Search-R1
Weimeng Luo
cs.IR, cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 16 pages, 3 figures. Code: https://github.com/luobostorm/search-r1-s2g-stopping
Code: https://github.com/luobostorm/search-r1-s2g-stopping
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper addresses the stopping problem in multi-round retrieval-augmented generation (RAG): deciding when to stop searching as evidence accumulates.
Terminology
Summary
This paper addresses the stopping problem in multi-round retrieval-augmented generation (RAG): deciding when to stop searching as evidence accumulates. The authors adapt S2G-RAG's structured sufficiency-and-gap judgment mechanism to a frozen Search-R1 pipeline and train a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions. The upstream system is the official Search-R1 Qwen2.5-7B PPO v0.2 checkpoint with E5 top-3 retrieval over Wiki-2018 and at most four searches.
The reasoner, retriever, corpus, prompt, and search budget remain unchanged; only the judge checkpoint and stopping threshold are selected on grouped validation and frozen before confirmatory evaluation.
The paper frames stopping as a sequential selection problem rather than an independent state-classification task
because the deployed policy is determined by the first STOP on each trajectory.
The evaluation chain has four layers: reachability (can the current state answer correctly now), state ranking (can the judge rank safe stop states), first crossing (which state crosses the threshold first), and system effects (answer quality, stop risk, and retrieval cost).
On an exploratory set of 200 questions with 723 native reachable states, the authors find that 143 of 523 native SEARCH actions (27.3%) occur at states where the frozen probe already produces the correct official answer.
These opportunities occur in 70 of 200 questions. The quality oracle reaches 51.5% EM versus native 46.5%, an exact gain of 10/200 or 5.0 percentage points.
The native-preserving cost oracle removes 131 of 523 searches (25.05%) while holding native EM fixed.
Fixed budgets do not realize the same trade-off: answer-only EM at budgets k = 0, 1, 2, 3, 4 is 18.5%, 34.0%, 43.0%, 46.0%, and 47.0%.
The exploratory mechanism ladder reveals that a binary LoRA raises STOP AP from 0.732 to 0.828 on the Core200 diagnostic, yet the zero-threshold policy reduces official EM from 36.5% to 31.5% and mean searches from 2.635 to 1.430.
The model learns a stronger state ranking signal but stops too aggressively. Replacing the binary target with a minimal S2G target reduces wrong STOP decisions from 99 to 24 and premature STOP decisions from 18 to 4,
with EM recovering to 35.5% and mean searches increasing to 2.390.
The judge is Qwen3.5-2B, trained for three epochs (12,654 steps) on 900 training questions with 3,009 clean states. The judge outputs a frozen JSON schema with a Boolean sufficient field and a list of gap items.
The online policy uses the log-probability margin for the Boolean field at a frozen prefix.
Grouped validation (100 questions) selects the threshold that maximizes recall subject to empirical STOP precision of at least 0.90 and at least 10 predicted STOP states.
The expanded judge freezes at a margin of approximately 7.875. Structured Base and the old LoRA use sentinels of 2.5 and approximately 8.0, respectively, because they cannot satisfy the validation precision-and-support rule.
On the confirmatory test set (indices 200–999 of the 1,000-question HotpotQA set), both predeclared criteria pass:
-
Official EM: Native reaches 0.44875; Expanded S2G reaches 0.44250. The paired difference is −0.00625 (95% CI [−0.01250, 0]), meaning a loss of 0.625 percentage points. The lower endpoint of the interval exceeds the frozen non-inferiority margin of −0.02 (two percentage points).
-
Mean searches: Native makes 2.60125 searches per question; Expanded S2G makes 2.50500. The difference is −0.09625 (95% CI [−0.12000, −0.07375]), with the upper endpoint below zero.
In absolute terms, Native makes 2,081 retrieval calls and the expanded policy makes 2,004, a reduction of 77 calls or 3.70%.
Six questions change from Native-correct to policy-wrong, while one changes from Native-wrong to policy-correct, yielding the net −5/800 = −0.625 percentage-point EM difference.
Expanded S2G improves STOP AP over Structured Base by 0.03983 (95% CI [0.00571, 0.07228]), providing confirmatory evidence of better state ranking. However, the policy is not safe: The expanded policy makes 69 system-level early stops, or 8.63% of the 800 questions. Forty-two are safe under the frozen target and 27 are unsafe, so the selected early-stop unsafe fraction is 39.13%.
The 42 safe selections cover only 14.29% of the 294 questions that contain a safe opportunity. State-level STOP precision is 0.6216 and recall is 0.0880. The grouped-validation STOP precision of 0.9091 is based on only 11 predicted STOP states and falls to 0.6216 at the frozen threshold on the confirmatory test set,
meaning the validation precision constraint does not transfer.
Calibration is weak: Expanded S2G has Brier score 0.2952 and ten-bin expected calibration error (ECE) 0.2911, compared with 0.1992 and 0.1179 for Structured Base.
At a post-hoc state-level operating point reaching at least 0.90 precision, recall falls to 0.0057.
The paper explicitly states: The result does not imply unchanged or improved accuracy, safe stopping, or lower total inference cost.
The expanded judge does not improve answer quality. Its point estimate is 0.625 percentage points below native EM.
The policy is also not a safe stopping rule
with 27 of 69 early stops unsafe. Finally, fewer retrieval calls do not prove lower total cost
because the Qwen3.5-2B judge's compute, memory, and latency are not converted into retrieval-equivalent units.
The paper concludes: "On the confirmatory test set, the complete policy makes 77 fewer retrieval calls than Native Search-R1: mean searches fall from 2.60125 to 2.50500, a difference of −0.09625 or 3.70% (95% CI [−0.12000, −0.07375]). Official EM falls from 0.44875 to 0.44250, a loss of 0.625 percentage points (95% CI [−1.25, 0] percentage points), which remains within the maximum acceptable loss of two percentage points specified before evaluation. The experiment therefore shows that the trained S2G-style judge reduces retrieval calls while broadly preserving answer accuracy. Here, 'broadly preserving' allows a small loss within the prespecified range; it does not mean that accuracy is unchanged or improved. The policy is not yet a safe stopping rule:
27 of its 69 early stops are unsafe, for selected-stop risk of 39.13%."
Improvements for AI systems
Improvements to AI Systems:
-
Add a calibrated early-stop risk gate: The judge’s STOP decisions are poorly calibrated (ECE 0.2911). Improve by training a secondary risk predictor that estimates the probability the current state’s answer is wrong, and only allow STOP when that probability is below a dynamic threshold (e.g., 0.10). This reduces the 39.13% unsafe early-stop fraction.
-
Replace fixed threshold with sequential hypothesis testing: Instead of a single margin threshold, use a sequential probability ratio test (SPRT) on the judge’s STOP log-probability margin across consecutive states. This would delay stopping when evidence is ambiguous, preventing premature stops (which caused 27 unsafe early stops) while still cutting retrieval calls.
-
Incorporate answer-consistency checks into the judge: Before stopping, require the judge to verify that the current answer is consistent with at least two independent retrieved passages (e.g., via entailment or token-overlap). This filters out spurious confident states, improving state-level STOP precision from 0.62 toward the 0.90 target.
-
Use a retrieval-cost-aware objective during judge training: The current judge optimizes sufficiency/gap labels, not retrieval cost. Retrain with a reward that penalizes unsafe stops more heavily than unnecessary searches, balancing the 3.70% call reduction against the 0.625% EM loss. This yields a policy that stops only when the expected answer-quality loss is negligible.
-
Add a fallback mechanism for low-confidence STOP states: When the judge’s margin is near the threshold (e.g., within 10% of the frozen margin), force an additional search instead of stopping. This would have avoided most of the 27 unsafe early stops while retaining most of the 77 saved retrieval calls.
-
Implement per-question adaptive search budgets: Use the judge’s sufficiency score to dynamically cap searches (e.g., stop early only if sufficiency > 0.9 and answer confidence > 0.8; otherwise, allow up to 4 searches). This prevents the aggressive early stopping seen with the zero-threshold LoRA and improves robustness across question difficulty.
What the Improved AI System Can Do:
-
Reduce retrieval calls by 3–5% while keeping EM loss under 0.5 percentage points (vs. the current 0.625% loss).
-
Achieve early-stop safety above 90% (i.e., fewer than 10% of early stops are wrong), compared to the current 39.13% unsafe rate.
-
Provide calibrated confidence scores for stopping decisions, enabling reliable human oversight and downstream task integration.
-
Adapt stopping behavior per question: safe to stop early on simple questions, but conservatively search more on ambiguous ones—improving both efficiency and accuracy over the current fixed-threshold policy.
Abstract
Multi-round retrieval-augmented generation (RAG) must decide when to stop searching as evidence accumulates. Because the deployed policy is determined by the first STOP on each trajectory, this is a sequential selection problem rather than an independent state-classification task. We adapt S2G-RAG's structured sufficiency-and-gap judgment to a frozen Search-R1 pipeline and train a Qwen3.5-2B judge on 3,009 states from 900 disjoint HotpotQA questions. Search-R1's reasoner, retriever, corpus, prompt, and search budget remain unchanged, while the judge checkpoint and stopping threshold are selected on grouped validation and frozen before confirmatory evaluation. On the confirmatory test set, the resulting policy reduces retrieval calls by 77 (3.70%) relative to Native Search-R1, while Official Exact Match decreases by 0.625 percentage points. Thus, the trained S2G-style structured judge reduces retrieval while broadly preserving answer accuracy. The result does not imply unchanged or improved accuracy, safe stopping, or lower total inference cost.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
- TASR: Training-Free Adaptive Stopping for Iterative Retrieval
- Formal Verification of Secure Encrypted Virtualization
- Selective Classification for Deep Neural Networks
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG