Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
cs.AI
Submitted: 2026-08-25
Updated: 2026-08-27
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/kookhh0827/copy-inflation-search-agents
License: http://creativecommons.org/licenses/by/4.0/
The gist: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning.
Terminology
Abstract
Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as token log probabilities, and has been actively studied for single-turn reasoning. However, modern LLMs increasingly act as multi-turn search agents that retrieve and condition on external documents. In this paper, we show that confidence-based voting transfers poorly to this multi-turn setting, and identify the underlying failure reason as copy inflation: when retrieved documents are appended to an agent's context, tokens copied from those documents receive systematically inflated log probabilities. This flattens confidence scores within each question and weakens the resulting weighted vote. To address this issue, we propose Retrieval-Grounded Voting (RGV), which scores each rollout by the lexical overlap between its final answer and the documents it retrieved. By computing the signal outside the contaminated context, RGV sidesteps both token log probabilities and additional LLM calls. Across four search-agent benchmarks and five LLMs, RGV consistently outperforms confidence-based voting, with gains of up to +5.4% accuracy and +35% on minority-correct questions, where the correct answer appears in only 1-2 of 8 rollouts.
Sources
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Universal Self-Consistency for Large Language Model Generation
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Training Verifiers to Solve Math Word Problems
- Deep Think with Confidence
- Retrieval-Augmented Generation for Large Language Models: A Survey
- GLM-5: from Vibe Coding to Agentic Engineering
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Language Models (Mostly) Know What They Know
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- Let's Verify Step by Step
- GAIA: a benchmark for General AI Assistants
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- WebGPT: Browser-assisted question-answering with human feedback
- gpt-oss-120b & gpt-oss-20b Model Card
- BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection