EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval

arXiv:2608.13006 · cs.CL, cs.IR · Submitted 2026-08-13 · Read on arXiv

Xinlong Xu, Yoshua Y. Li

Nanjing University of Information Science and Technology · Meituan

cs.CL, cs.IR

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/XrazyMee/EviReform

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 95/100

The gist: EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval Summary This paper introduces EviReform, a method for multi-hop graph retrieval that separates revising the retrieval

Terminology

Summary

EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval

Summary

This paper introduces EviReform, a method for multi-hop graph retrieval that separates revising the retrieval request from aggregating evidence in the graph. The core idea is that an initially retrieved passage can resolve an entity or relation implicit in the question, making the remaining information need more specific. EviReform uses this observation to formulate residual queries for the unresolved need, combines their retrieval signals with the original question, and propagates the result through shared entities.

Problem and Motivation

The paper argues that structure alone does not capture every change introduced by retrieval. Suppose an initial passage identifies the person, location, or relation implicit in the question. That observation makes the remaining information need more specific than it was before retrieval. A graph retriever whose seeds, paths, or edge scores remain tied to the original question must recover the complementary passage through stored relations, even though the observed passage provides a more direct description of what is missing.

The paper distinguishes two decisions: before any passage is read, the question can determine where retrieval enters the graph and how it is traversed; after a passage is read, the system can reconsider what textual evidence it should seek. "Better traversal does not remove the need for this second decision when the bridge is implicit, absent from the graph, or more directly expressed in passage text. Conversely, reformulation does not replace graph structure: the evidence retrieved for the revised query can still be incomplete or dispersed across related propositions."

Method

EviReform implements this separation directly. The method works as follows:

  1. Initial Evidence Selection: Given a question q with normalized embedding hq, EviReform scores propositions by ui(q) = max(0, h⊤i hq). An LLM selects a set Sq of proposition identifiers from the candidates with the highest scores. The complete source passages of the selected propositions form the observed evidence Eq = dπ(i): i ∈ Sq.

  2. Query Reformulation: EviReform passes (q, Eq) to an LLM that produces residual queries ρ = ρ1,..., ρL. Each residual query describes information needed for the original question but not established by Eq. A useful residual query preserves the constraints of the original question, incorporates a bridge established by Eq, and asks for the unresolved relation or attribute.

  3. Signal Combination: For each residual query ρl, dense retrieval selects a set Il of propositions. Its normalized signal is r(l)i = I[i ∈ Il] max(0, h⊤i hρl) / Σj∈Il max(0, h⊤j hρl). The signals are averaged and combined with the seed from the original question: s = βb + (1 − β)r, where s = b if no residual query is valid. Independent normalization gives the original and residual channels controlled total mass.

  4. Propagation through Shared Entities: The combined signal is propagated between propositions that share entities. The proposition weights are W = AD†e A⊤ − diag(AD†e A⊤), with transition T = WD†w. One update is applied: z = αs + (1 − α)Ts, where α balances direct retrieval evidence against support transferred through shared entities.

  5. Passage Readout: Proposition scores are aggregated to passages with score(dj) = (1/√Pj) Σi:π(i)=j zi, where Pj is the set of propositions extracted from dj.

Key Design Choices

  • The reformulator reads passage text rather than isolated proposition strings, supplying surrounding facts needed to identify what remains unresolved.

  • Initial selection is deliberately not treated as the final retrieval result; selected passages must earn their final positions through the same combined scoring and propagation used for all other passages.

  • The reformulator produces retrieval queries rather than an answer or reasoning trace, and the system returns one passage ranking.

  • Propagation occurs after the two retrieval signals are combined, consolidating propositions reached from either signal when they share entities.

Experimental Setup

The method is evaluated on 2WikiMultiHopQA, HotpotQA, MuSiQue, and GraphRAG-Bench (Medical). The corpora contain 6,119, 9,811, and 11,656 passages, respectively. Baselines include BM25, BGE-M3, Qwen3-Embedding-0.6B, dense with rerankers, GritLM-7B, NV-Embed-v2, IRCoT, S2G-RAG, GeAR, PropRAG, HippoRAG 2, and CatRAG. EviReform uses 100 initial proposition candidates, at most 12 selected propositions, at most three residual queries with two propositions retrieved per query, and α = β = 0.5.

Main Results

Relative to the strongest passage ranker for each metric, EviReform gains 5.00, 2.65, and 5.59 R@5 points on 2Wiki, HotpotQA, and MuSiQue; the R@10 gains are 0.88, 1.20, and 4.78 points. With the shared reader, EviReform improves over the strongest QA baseline, GeAR, by 3.91, 2.28, and 4.50 F1 points, and by 3.30, 1.60, and 4.00 EM points.

Paired confidence intervals use 10,000 question-level bootstrap resamples. The R@5 intervals are [3.90, 6.10], [1.50, 3.80], and [3.86, 7.28] points; the F1 intervals are [1.42, 6.36], [0.34, 4.24], and [2.31, 6.72]. Among the remaining metrics, only the HotpotQA EM interval overlaps zero.

EviReform improves Chain@5 by 22.5, 12.1, and 11.7 points over the strongest graph baselines. On GraphRAG-Bench (Medical), EviReform reaches 71.75 mean answer accuracy, compared with 67.48 for S2G-RAG, 69.25 for GeAR, and 69.86 for HippoRAG 2.

Mechanism Analysis

The 2×2 ablation crosses query reformulation with propagation. All three datasets satisfy Full > Reformulation only > Propagation only > Base on R@5 and Chain@5. With propagation present, reformulation adds 7.20/18.00, 2.75/5.40, and 3.63/5.40 R@5/Chain@5 points. After reformulation, propagation adds a further 0.60/1.20, 0.60/1.20, and 0.93/2.10 points.

A matched retrieval run that repeats the original question instead of using residual queries reaches 89.75/74.70, 93.90/88.20, and 67.51/39.20 R@5/Chain@5, compared with 97.75/94.90, 96.70/93.80, and 73.03/46.90 for the full method. The gain therefore comes from specifying the unresolved need, rather than issuing more requests for the original question.

Reranking the initial candidates with first-stage evidence reaches 73.83/50.10, 88.75/80.90, and 59.48/29.90 R@5/Chain@5. These pools contain a complete chain for only 53.0%, 90.2%, and 43.9% of questions.

Pool coverage analysis shows that before reformulation, initial pools cover 78.8%, 95.3%, and 74.8% of individual gold passages, and contain complete chains for 54.1%, 90.7%, and 46.7% of questions. After residual retrieval, passage coverage rises to 98.1%, 98.4%, and 83.0%, while complete-chain coverage rises to 95.7%, 96.9%, and 61.7%.

Additional Analyses

  • With NV-Embed-v2 used throughout the graph retrievers, EviReform exceeds the strongest graph baseline by 7.20, 1.00, and 5.17 R@5 points on 2Wiki, HotpotQA, and MuSiQue.

  • Allowing up to three rounds of reformulation changes R@5 by −0.20, −0.40, and +2.47 points on 2Wiki, HotpotQA, and MuSiQue relative to one round. Only MuSiQue receives a substantial gain, so one round is used for all three datasets.

  • The response weight α is comparatively stable: its R@5 range is 0.12 points on 2Wiki, 0.50 on HotpotQA, and 1.45 on MuSiQue. The mass β assigned to the original question has a wider range, especially on MuSiQue, where increasing it from 0.25 to 0.75 raises R@5 from 70.47 to 74.18 and Chain@5 from 44.80 to 49.10.

  • As the index grows from 3,254 to 11,656 passages on MuSiQue, R@5 changes from 76.46 to 72.58, R@10 from 86.96 to 80.54, and R@20 from 89.92 to 85.50.

Conclusion

"Multi-hop retrieval changes as evidence is acquired: an initial passage can reveal the entity or relation needed to find its complement. EviReform uses that observation to formulate residual queries, combines their retrieval signal with the original question, and propagates the result between propositions that share entities. Across three multi-hop benchmarks, this design improves passage recall, complete-chain recovery, and downstream QA. Mechanism studies show that reformulation recovers most of the new evidence, while propagation consolidates evidence reached by both requests. EviReform thus lets observed evidence refine what graph retrieval searches for before producing the final passage ranking."

Improvements for AI systems

Improvements to AI Systems:

  1. Residual Query Generation for Multi-Step Reasoning: Implement a module that, after retrieving initial evidence, prompts the LLM to generate residual queries—specific, unresolved information needs that incorporate newly discovered entities or relations. This transforms single-pass retrieval into an iterative, evidence-aware process, enabling the system to ask what is still missing? rather than re-issuing the original question.

  2. Dual-Channel Retrieval Signal Fusion: Combine the original query's dense retrieval signal with the residual queries' signals using normalized, weighted averaging (e.g., β for original, 1−β for residuals). This prevents the residual channel from dominating while ensuring both evidence sources contribute proportionally, improving recall of complementary passages.

  3. Entity-Shared Propagation for Evidence Consolidation: After fusing retrieval signals, propagate scores between propositions that share entities using a graph-based transition matrix. This consolidates evidence reached via different queries, filling gaps where a single query fails to connect all necessary facts, and improves complete-chain recovery.

  4. Passage-Level Aggregation with Normalization: Aggregate proposition-level scores to passage-level using inverse-square-root normalization (1/√P j) to avoid bias toward passages with many propositions. This yields fairer passage rankings, especially in dense corpora.

  5. Controlled Reformulation Rounds: Limit reformulation to one round by default, with an option to extend to three rounds only when substantial gains are observed (e.g., MuSiQue). This prevents over-iteration costs while capturing the primary benefit of evidence-guided refinement.

  6. Adaptive Weight Tuning for Query Mass: Dynamically adjust the weight β assigned to the original question based on dataset complexity. Higher β (e.g., 0.75) improves performance on harder multi-hop datasets, suggesting the system should prioritize original constraints when residual queries are less reliable.

Capabilities of the Improved AI System:

  • Iterative Evidence Refinement: The system can now refine its search based on what it has already read, resolving implicit entities or relations and generating targeted follow-up queries—enabling it to handle questions where the bridge between facts is not explicitly in the graph.

  • Robust Multi-Hop Retrieval: It achieves significantly higher passage recall (e.g., +5.59 R@5 on MuSiQue) and complete-chain recovery (e.g., +22.5 Chain@5 on 2Wiki) compared to state-of-the-art graph retrievers, by combining original and residual signals with entity-based propagation.

  • Better Downstream QA Performance: With a shared reader, the system improves answer F1 by up to 4.50 points and EM by up to 4.00 points over strong baselines like GeAR, demonstrating that better retrieval directly translates to more accurate answers.

  • Scalable and Adaptive: It maintains performance across corpus sizes (from 3K to 11K passages) and adapts its reformulation strategy per dataset, ensuring efficiency without sacrificing accuracy.

  • Transparent Evidence Consolidation: The system can explain its final ranking by showing how initial evidence, residual queries, and entity-sharing propagation contributed to each passage's score, aiding interpretability in complex reasoning tasks.

Abstract

Multi-hop retrieval must recover passages that provide sufficient evidence together. An initial passage often resolves an entity or relation implicit in the question, making the missing evidence easier to describe only after retrieval begins. Graph retrieval improves access to related evidence through stored corpus structure, but its retrieval signal is commonly derived from the original question. Complementary evidence must then be reached through stored relations even when an observed passage provides a more direct semantic cue. We introduce EviReform, which separates revising the retrieval request from aggregating evidence in the graph. Retrieved source passages formulate residual queries for the unresolved information need. The original and residual retrieval signals are normalized separately, combined, and propagated between propositions that share entities. On 2WikiMultiHopQA, HotpotQA, and MuSiQue, EviReform exceeds the strongest baseline by up to 5.59 Recall@5 points and 4.50 F1 points. These results show that observed evidence can guide graph retrieval toward the part of a supporting chain left underspecified by the original question. Code is available at https://github.com/XrazyMee/EviReform.

Sources

Related papers