ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

arXiv:2608.10928 · cs.AI · Submitted 2026-08-11 · Read on arXiv

Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha

IIT Bombay · University of Maryland, College Park · Adobe Research

cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/itsvaibhav01/ThinkRetrieve

Project page: https://itsvaibhav01.github.io/ThinkRetrieve

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: ThinkRetrieve is a test-time scaling framework that augments the reasoning traces of Large Reasoning Models (LRMs) with dynamically retrieved solved examples at each reasoning step.

Terminology

Summary

ThinkRetrieve is a test-time scaling framework that augments the reasoning traces of Large Reasoning Models (LRMs) with dynamically retrieved solved examples at each reasoning step. The paper states: "We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant."

The core idea is that after each reasoning step, the model generates an intermediate answer, which is used as a query to retrieve a relevant solved example from an external example bank. The paper explains: "After the initial thinking phase, the model generates an intermediate answer, which is used as a query to retrieve a relevant solved example from an external example bank. The retrieved exemplar is inserted into the thinking trace as an in-context example (ICE), enabling the model to extract key takeaways, identify errors in its own reasoning, and course-correct before producing the final answer."

This approach contrasts with standard sequential test-time scaling, which the paper notes often yields diminishing or negative returns: "recent studies show that thinking longer does not always mean thinking better... as traces grow, LRMs exhibit increased uncertainty, repetitive cycling, drift, and error compounding, and additional compute often amplifies rather than corrects the mistake. The paper draws an analogy to human problem-solving: This contrasts with how humans tackle hard problems: instead of persisting along a single line of thought, we recall analogous solved problems to guide and verify our reasoning."

The framework's methodology is formalized as an interleaved reasoning trajectory: "ThinkRetrieve proceeds iteratively, constructing an interleaved reasoning trajectory τk = (z1, e1, z2, e2,..., zk, ek), where zt denotes the reasoning step generated by the model at step t, et denotes the exemplar retrieved after zt, and k is the total number of reasoning steps determined by the token budget B." At each step, the model generates a reasoning trace, produces an intermediate answer via the thinking delimiter token, and this intermediate answer is jointly encoded with the test query to retrieve the most relevant exemplar via dense nearest-neighbor search.

The paper reports extensive experiments across five reasoning models (DeepSeek-R1-Distill-Qwen-1.5B, Qwen3-1.7B, Qwen3.5-2B, Qwen3-4B, and Qwen3-8B) on four benchmarks (GSM-8K, MATH-500, AIME 2025, and SciQ). The key result is: ThinkRetrieve wins on every (model, benchmark) cell across five reasoning models and four benchmarks... with an absolute gain of up to +13.4 points on AIME 2025. The paper highlights that "ThinkRetrieve maintains monotonically increasing or stable accuracy as the thinking budget grows, demonstrating that retrieval-augmented test-time scaling uses additional compute more effectively than self-reflection alone."

The paper provides detailed analysis of why ThinkRetrieve helps. It shows that retrieved exemplars reduce the model's uncertainty over its final answer at each reasoning step, preventing the error accumulation that causes reasoning drift under sequential self-reflection. This is supported by two measures: predictive entropy (ThinkRetrieve lowers entropy by approximately 0.55 nats) and per-step confidence (under ThinkRetrieve, the length-normalized negative log-likelihood of the final answer decreases monotonically; each retrieved exemplar anchors the model's belief and prevents the late-stage confidence reversal).

The paper also includes extensive ablations and controls. It shows that both per-step in-trace injection and semantic retrieval relevance are necessary: "S-ICL (static input-level ICL) prepends k=3 QA-conditioned retrieved exemplars to the prompt once before reasoning begins... Rand (random retrieval) uses the same per-step injection mechanism as ThinkRetrieve but selects uniformly random corpus exemplars... Both baselines underperform ThinkRetrieve on every (model, benchmark) cell. A compute-matched comparison against self-consistency shows that single-call ThinkRetrieve outperforms every TTS self-consistency configuration at k ∈ 2, 4, 8 by 16–27 absolute accuracy points," ruling out the hypothesis that gains come from compute alone.

The paper addresses potential leakage concerns through a two-stage decontamination process and multiple audits. The leakage audit confirms Zero queries exceed the removal threshold of 0.90, and the maximum observed retained similarity is 0.898. An answer-leakage control with an additional retrieval-time filter excluding corpus entries matching the gold answer shows Accuracy is preserved — in fact it slightly improves on a controlled evaluation subset — confirming that the method's gain reflects structural rather than answer-level similarity. The top-1 retrieved-exemplar audit confirms Zero of the audited exemplars share their boxed final answer with the test problem they were retrieved for.

The paper also provides qualitative examples showing ThinkRetrieve's effectiveness. In one MATH-500 probability problem, both methods arrive at the same incorrect intermediate answer of 1/6, but "ThinkRetrieve retrieves a structurally similar solved problem whose solution highlights a key counting distinction the model had overlooked, prompting it to correct 1/6 to 2/6 and arrive at the correct final answer of 1/3. In another geometry problem, sequential TTS second-guesses its correct solution, explores alternative parametric and rotational proof strategies, encounters algebraic errors, and ultimately produces an incorrect final answer, while ThinkRetrieve retrieves a related problem that anchors the model's confidence in its existing solution."

The paper acknowledges several limitations, including dependence on corpus coverage, potential for misleading exemplars, additional latency from retrieval calls, and the use of domain-matched corpora. It notes: "for problems distributionally distant from E—whether in domain, difficulty, or required reasoning style—retrieved exemplars may be irrelevant or actively misleading, potentially degrading performance below the no-retrieval baseline. The paper concludes that ThinkRetrieve consistently improves accuracy over sequential test-time scaling, maintains monotonically increasing performance as the thinking budget grows, and reduces answer entropy across reasoning steps."

Improvements for AI systems

Improvements to AI Systems:

  1. Dynamic In-Trace Exemplar Injection: Implement a retrieval mechanism that operates during the reasoning process, not just at input. After each intermediate reasoning step, the system generates a query from the current partial answer, retrieves a structurally similar solved problem from an external corpus, and injects it directly into the thinking trace. This allows the model to course-correct mid-reasoning rather than committing to a single flawed trajectory.

  2. Semantic Retrieval with Joint Query Encoding: Use dense nearest-neighbor search where the test query and the current intermediate answer are jointly encoded to retrieve exemplars. This ensures the retrieved example is relevant to the current reasoning state, not just the original problem, enabling the model to identify and fix specific logical errors as they emerge.

  3. Confidence Anchoring via Retrieved Exemplars: Integrate retrieved exemplars to reduce predictive entropy and stabilize per-step confidence. The system should use these exemplars to anchor the model's belief in its final answer, preventing late-stage confidence reversal and error compounding that occurs in long sequential self-reflection.

  4. Monotonic Budget Scaling: Design the system so that increasing the token budget for thinking leads to monotonically increasing or stable accuracy, unlike standard sequential scaling which shows diminishing or negative returns. Each retrieved exemplar should serve as a corrective checkpoint, ensuring additional compute translates to better reasoning rather than more repetitive cycling or drift.

  5. Structural Similarity over Answer Leakage: Implement a retrieval filter that excludes exemplars sharing the gold answer with the test problem, forcing the system to rely on structural and methodological similarity. This prevents the model from copying answers and instead teaches it how to reason through analogous problem-solving steps.

  6. Error-Correction Trigger Mechanism: Use the retrieved exemplar to explicitly prompt the model to extract key takeaways and compare them against its own reasoning. The system should be able to identify discrepancies between its intermediate answer and the exemplar's solution path, then generate a corrected reasoning step before producing the final answer.

What the Improved AI System Can Do:

  • Solve complex mathematical and scientific reasoning problems (e.g., AIME 2025, MATH-500) with higher accuracy, achieving up to +13.4 absolute points over baseline sequential reasoning.

  • Maintain consistent performance gains across multiple model sizes (1.5B to 8B parameters) and diverse benchmarks (GSM-8K, SciQ, etc.), demonstrating scalability and robustness.

  • Avoid reasoning drift and error accumulation in long thinking traces by anchoring each step with relevant solved examples, reducing final-answer uncertainty by 0.55 nats.

  • Outperform self-consistency methods (e.g., majority voting over multiple samples) by 16–27 absolute accuracy points using a single call, making it more compute-efficient.

  • Adapt to new problems by retrieving and applying analogous solutions in real-time, mimicking human recall of similar past problems to guide verification and correction.

  • Provide reliable performance even when the external corpus is imperfect, as long as the retrieved exemplars are structurally relevant, with safeguards against misleading or answer-leaking examples.

Sources

Related papers