ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha
IIT Bombay · University of Maryland, College Park · Adobe Research
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/itsvaibhav01/ThinkRetrieve
Project page: https://itsvaibhav01.github.io/ThinkRetrieve
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: ThinkRetrieve is a test-time scaling framework that augments the reasoning traces of Large Reasoning Models (LRMs) with dynamically retrieved solved examples at each reasoning step.
Terminology
Summary
ThinkRetrieve is a test-time scaling framework that augments the reasoning traces of Large Reasoning Models (LRMs) with dynamically retrieved solved examples at each reasoning step. The paper states: "We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant."
The core idea is that after each reasoning step, the model generates an intermediate answer, which is used as a query to retrieve a relevant solved example from an external example bank. The paper explains: "After the initial thinking phase, the model generates an intermediate answer, which is used as a query to retrieve a relevant solved example from an external example bank. The retrieved exemplar is inserted into the thinking trace as an in-context example (ICE), enabling the model to extract key takeaways, identify errors in its own reasoning, and course-correct before producing the final answer."
This approach contrasts with standard sequential test-time scaling, which the paper notes often yields diminishing or negative returns: "recent studies show that thinking longer does not always mean thinking better... as traces grow, LRMs exhibit increased uncertainty, repetitive cycling, drift, and error compounding, and additional compute often amplifies rather than corrects the mistake. The paper draws an analogy to human problem-solving:
This contrasts with how humans tackle hard problems: instead of persisting along a single line of thought, we recall analogous solved problems to guide and verify our reasoning."
The framework's methodology is formalized as an interleaved reasoning trajectory: "ThinkRetrieve proceeds iteratively, constructing an interleaved reasoning trajectory τk = (z1, e1, z2, e2,..., zk, ek), where zt denotes the reasoning step generated by the model at step t, et denotes the exemplar retrieved after zt, and k is the total number of reasoning steps determined by the token budget B." At each step, the model generates a reasoning trace, produces an intermediate answer via the thinking delimiter token, and this intermediate answer is jointly encoded with the test query to retrieve the most relevant exemplar via dense nearest-neighbor search.
The paper reports extensive experiments across five reasoning models (DeepSeek-R1-Distill-Qwen-1.5B, Qwen3-1.7B, Qwen3.5-2B, Qwen3-4B, and Qwen3-8B) on four benchmarks (GSM-8K, MATH-500, AIME 2025, and SciQ). The key result is: ThinkRetrieve wins on every (model, benchmark) cell across five reasoning models and four benchmarks... with an absolute gain of up to +13.4 points on AIME 2025.
The paper highlights that "ThinkRetrieve maintains monotonically increasing or stable accuracy as the thinking budget grows, demonstrating that retrieval-augmented test-time scaling uses additional compute more effectively than self-reflection alone."
The paper provides detailed analysis of why ThinkRetrieve helps. It shows that retrieved exemplars reduce the model's uncertainty over its final answer at each reasoning step, preventing the error accumulation that causes reasoning drift under sequential self-reflection.
This is supported by two measures: predictive entropy (ThinkRetrieve lowers entropy by approximately 0.55 nats) and per-step confidence (under ThinkRetrieve, the length-normalized negative log-likelihood of the final answer decreases monotonically; each retrieved exemplar anchors the model's belief and prevents the late-stage confidence reversal
).
The paper also includes extensive ablations and controls. It shows that both per-step in-trace injection and semantic retrieval relevance are necessary: "S-ICL (static input-level ICL) prepends k=3 QA-conditioned retrieved exemplars to the prompt once before reasoning begins... Rand (random retrieval) uses the same per-step injection mechanism as ThinkRetrieve but selects uniformly random corpus exemplars... Both baselines underperform ThinkRetrieve on every (model, benchmark) cell. A compute-matched comparison against self-consistency shows that
single-call ThinkRetrieve outperforms every TTS self-consistency configuration at k ∈ 2, 4, 8 by 16–27 absolute accuracy points," ruling out the hypothesis that gains come from compute alone.
The paper addresses potential leakage concerns through a two-stage decontamination process and multiple audits. The leakage audit confirms Zero queries exceed the removal threshold of 0.90, and the maximum observed retained similarity is 0.898.
An answer-leakage control with an additional retrieval-time filter excluding corpus entries matching the gold answer shows Accuracy is preserved — in fact it slightly improves on a controlled evaluation subset — confirming that the method's gain reflects structural rather than answer-level similarity.
The top-1 retrieved-exemplar audit confirms Zero of the audited exemplars share their boxed final answer with the test problem they were retrieved for.
The paper also provides qualitative examples showing ThinkRetrieve's effectiveness. In one MATH-500 probability problem, both methods arrive at the same incorrect intermediate answer of 1/6, but "ThinkRetrieve retrieves a structurally similar solved problem whose solution highlights a key counting distinction the model had overlooked, prompting it to correct 1/6 to 2/6 and arrive at the correct final answer of 1/3. In another geometry problem, sequential TTS
second-guesses its correct solution, explores alternative parametric and rotational proof strategies, encounters algebraic errors, and ultimately produces an incorrect final answer, while ThinkRetrieve retrieves a related problem that
anchors the model's confidence in its existing solution."
The paper acknowledges several limitations, including dependence on corpus coverage, potential for misleading exemplars, additional latency from retrieval calls, and the use of domain-matched corpora. It notes: "for problems distributionally distant from E—whether in domain, difficulty, or required reasoning style—retrieved exemplars may be irrelevant or actively misleading, potentially degrading performance below the no-retrieval baseline. The paper concludes that
ThinkRetrieve consistently improves accuracy over sequential test-time scaling, maintains monotonically increasing performance as the thinking budget grows, and reduces answer entropy across reasoning steps."
Improvements for AI systems
Improvements to AI Systems:
-
Dynamic In-Trace Exemplar Injection: Implement a retrieval mechanism that operates during the reasoning process, not just at input. After each intermediate reasoning step, the system generates a query from the current partial answer, retrieves a structurally similar solved problem from an external corpus, and injects it directly into the thinking trace. This allows the model to course-correct mid-reasoning rather than committing to a single flawed trajectory.
-
Semantic Retrieval with Joint Query Encoding: Use dense nearest-neighbor search where the test query and the current intermediate answer are jointly encoded to retrieve exemplars. This ensures the retrieved example is relevant to the current reasoning state, not just the original problem, enabling the model to identify and fix specific logical errors as they emerge.
-
Confidence Anchoring via Retrieved Exemplars: Integrate retrieved exemplars to reduce predictive entropy and stabilize per-step confidence. The system should use these exemplars to anchor the model's belief in its final answer, preventing late-stage confidence reversal and error compounding that occurs in long sequential self-reflection.
-
Monotonic Budget Scaling: Design the system so that increasing the token budget for thinking leads to monotonically increasing or stable accuracy, unlike standard sequential scaling which shows diminishing or negative returns. Each retrieved exemplar should serve as a corrective checkpoint, ensuring additional compute translates to better reasoning rather than more repetitive cycling or drift.
-
Structural Similarity over Answer Leakage: Implement a retrieval filter that excludes exemplars sharing the gold answer with the test problem, forcing the system to rely on structural and methodological similarity. This prevents the model from copying answers and instead teaches it how to reason through analogous problem-solving steps.
-
Error-Correction Trigger Mechanism: Use the retrieved exemplar to explicitly prompt the model to extract key takeaways and compare them against its own reasoning. The system should be able to identify discrepancies between its intermediate answer and the exemplar's solution path, then generate a corrected reasoning step before producing the final answer.
What the Improved AI System Can Do:
-
Solve complex mathematical and scientific reasoning problems (e.g., AIME 2025, MATH-500) with higher accuracy, achieving up to +13.4 absolute points over baseline sequential reasoning.
-
Maintain consistent performance gains across multiple model sizes (1.5B to 8B parameters) and diverse benchmarks (GSM-8K, SciQ, etc.), demonstrating scalability and robustness.
-
Avoid reasoning drift and error accumulation in long thinking traces by anchoring each step with relevant solved examples, reducing final-answer uncertainty by 0.55 nats.
-
Outperform self-consistency methods (e.g., majority voting over multiple samples) by 16–27 absolute accuracy points using a single call, making it more compute-efficient.
-
Adapt to new problems by retrieving and applying analogous solutions in real-time, mimicking human recall of similar past problems to guide verification and correction.
-
Provide reliable performance even when the external corpus is imperfect, as long as the retrieved exemplars are structurally relevant, with safeguards against misleading or answer-leaking examples.
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- Guideline Forest: Retrieval-Augmented Reasoning with Branching Experience-Induced Guidelines
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3 Technical Report
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Thinkless: LLM Learns When to Think
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Inverse Scaling in Test-Time Compute
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Think Only When You Need with Large Hybrid-Reasoning Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- ThinkSwitcher: When to Think Hard, When to Think Fast
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Dr.ICL: Demonstration-Retrieved In-context Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection