Calibrated Fusion for Heterogeneous Graph-Vector Retrieval in Multi-Hop QA
cs.IR, cs.LG
Submitted: 2026-03-30
Updated: 2026-09-18
Comments: v4: MuSiQue LastHop recomputed against the terminal-hop passage (v1-v3 scored the last supporting paragraph in MuSiQue's paragraph order): vector-only 69.1 -> PhaseGraph 71.0 at @10 (15W/5L, p=.041); True RRF 71.8 (25W/11L, p=.029); no paired advantage over RRF on either benchmark; embedding-realization sensitivity disclosed. 10 pages, 6 figures, 9 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Graph-augmented retrieval combines dense similarity with graph-based relevance signals such as Personalized PageRank (PPR), but these scores have different distributions and are not directly
Terminology
Abstract
Graph-augmented retrieval combines dense similarity with graph-based relevance signals such as Personalized PageRank (PPR), but these scores have different distributions and are not directly comparable. We study this as a score calibration problem for heterogeneous retrieval fusion in multi-hop question answering. Our method, PhaseGraph, maps vector and graph scores to a common unit-free scale using percentile-rank normalization (PIT) before fusion, enabling stable combination without discarding magnitude information. Across MuSiQue and 2WikiMultiHopQA, calibrated fusion improves held-out last-hop retrieval on HippoRAG2-style benchmarks: LastHop@10 increases from 69.1% to 71.0% on MuSiQue (15W/5L, p=0.041, n=514) and LastHop@5 from 51.7% to 53.6% on 2WikiMultiHopQA (11W/2L, p=0.023, n=491), both on independent held-out test splits. Against the official HippoRAG 2 pipeline on 2WikiMultiHopQA, calibrated fusion is ahead at LastHop@10 (+6.3pp, p<10-3) and behind at LastHop@5 (-8.4pp), a cutoff-dependent cross-over we report in full. A theory-driven ablation shows that percentile-based calibration is directionally more robust than min-max normalization on both tune and test splits (1W/6L, p=0.125), while Boltzmann weighting performs comparably to linear fusion after calibration (0W/3L, p=0.25). These results suggest that score commensuration is a robust design choice, and the exact post-calibration operator appears to matter less on these benchmarks.
Sources
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
- PropRAG: Guiding Retrieval with Beam Search over Proposition Paths
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG