SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges
Yuchao Wu, Junqin Li, XingCheng Liang, Yongjie Chen, Yinghao Liang, Linyuan Mo, Guanxian Li
Zleap AI
cs.CL
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/Zleap-AI/SAG-Benchmark
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges" 1.
Terminology
Summary
Summary of the Paper SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges
1. Core Problem and Motivation
The paper addresses the limitations of existing retrieval-augmented generation (RAG) systems for multi-hop question answering. It states that mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning.
While graph-based methods like GraphRAG address this by constructing knowledge graphs offline, they often fragment semantics, incur high maintenance, and complicate incremental updates.
The central research question posed is: can the structure needed by a query be activated from an append-only index during retrieval, without constructing a global knowledge graph?
2. Proposed Method: SAG
The authors propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph.
-
Event-Entity Index: Each chunk is processed independently into one semantically complete event and a set of entities. This pair defines a
latent hyperedge
thatpreserves n-ary relations without decomposing them into triples.
The index is append-only, allowing new documents to be addedwithout recomputing existing records.
-
Query-Time Dynamic Hyperedges: At query time, SAG
treats shared entities as join keys to connect related chunks
using SQL joins. Thisdynamically yields a query-scoped neighborhood of events,
andevery piece of evidence remains the original chunk throughout.
-
Architecture: The system has an offline indexing phase and an online phase with seed retrieval, query-time expansion, and final selection. It uses SQL for exact entity joins, vector retrieval for semantic similarity, and an LLM for event-entity extraction, query-entity identification, and final selection.
3. Key Technical Details
-
Seed Retrieval: SAG uses two parallel paths: Path A (entity-guided structured recall) identifies query entities and uses SQL joins to find events sharing those entities; Path B (direct event recall) retrieves events by embedding similarity to the query.
-
Expansion: Starting from seed events, SAG uses reverse SQL joins to find new entities and then new events, performing
event-to-entity-to-event composition
for up to L rounds (default L=1). -
Final Selection: An LLM reads the top-K candidates and selects up to 5 events that jointly entail the answer. This is a
contextual selection step rather than a pointwise reranking.
The semantic path fills the remaining slots to return 10 chunks total.
4. Experimental Setup
-
Datasets: MuSiQue (up to 4 hops), 2WikiMultiHopQA, and HotpotQA (2 hops). For corpus growth, they also use NQ.
-
Baselines: Compared against simple retrievers (BM25, Contriever), large embedding models (BGE-Large, GTE-Qwen2, GritLM, NV-Embed), and structure-augmented methods (GraphRAG, LightRAG, HippoRAG 2, HyperGraphRAG, HyperRAG).
-
Unified Configuration: All structure-augmented methods use BGE-Large-EN-v1.5 as the retriever and Qwen3.6-Flash as the reader to isolate architecture from model choice.
5. Main Results
SAG achieves the best retrieval and end-to-end QA performance on all three benchmarks, with gains widening as reasoning-chain complexity increases.
-
Retrieval (Recall@5):
-
On MuSiQue, SAG reaches 80.36%, outperforming the strongest baseline (HippoRAG 2 at 65.13%) by 15.23 points.
-
On 2WikiMultiHopQA, SAG reaches 93.34%, surpassing HippoRAG 2 (90.35%).
-
On HotpotQA, SAG reaches 96.50%, leading HippoRAG 2 (94.35%) by 2.15 points.
-
QA (F1):
-
On MuSiQue, SAG achieves 61.15 F1, exceeding GraphRAG by 7.01 points.
-
On 2WikiMultiHopQA and HotpotQA, SAG leads the strongest baseline by 1.17 and 1.66 points, respectively.
6. Analysis and Ablations
-
Ablation on MuSiQue: Final selection has the largest effect (replacing the LLM with a reranker reduces Recall@5 by 13.25 points). Disabling expansion reduces Recall@5 by 10.95 points. Replacing hyperedge indexing with triple indexing causes a smaller decrease of 2.75 points.
-
Embedding Robustness: SAG is more robust to embedding model changes. Replacing NV-Embed-v2 with BGE-v1.5 reduces HippoRAG 2 Recall@5 by 9.42 points, but SAG only drops by 1.35 points.
-
Scaling: With a candidate pool of 500 (out of 11,656), SAG retains 99% of the full-pool result (79.57% vs 80.36% Recall@5), showing bounded activation limits downstream workloads.
-
Corpus Growth: SAG degrades more slowly than HippoRAG 2 as the corpus grows. On MuSiQue, SAG drops from 87.90% to 82.57% (loss of 5.33 points), while HippoRAG 2 falls from 69.23% to 60.27% (loss of 8.96 points).
7. Limitations
The paper identifies two main limitations:
-
No Alias Resolution: SAG
normalizes and deduplicates entity strings but does not resolve aliases,
so it cannot recognize thatApple Inc.
andApple
refer to the same entity. -
No Temporal Updates: The index is append-only, providing
no mechanism for revising or retiring stale events,
which is insufficient for agent memory that must represent changing facts and preferences.
8. Conclusion
The paper concludes that SAG turns indexed data into structured context for multi-hop retrieval,
achieving the strongest average retrieval and QA performance under unified settings. The authors state that event representation, query-time expansion, and LLM-based final selection each contribute to retrieval quality,
and that SAG varies less across embedding models and degrades more slowly as the corpus grows.
Future work will explore alias resolution and versioned updates.
Improvements for AI systems
Improvements to AI Systems Based on SAG:
-
Dynamic Query-Time Structuring: Implement an append-only event-entity index where each document chunk is stored as a semantically complete event with linked entities. At query time, use SQL-style joins on shared entities to dynamically construct query-scoped hyperedges, eliminating the need for pre-built global knowledge graphs. This enables the system to handle multi-hop reasoning without costly offline graph construction or maintenance.
-
Hybrid Retrieval with Dual Pathways: Integrate two parallel retrieval paths—entity-guided structured recall (using exact SQL joins on query entities) and semantic embedding-based event recall. Combine results to capture both precise relational constraints and fuzzy semantic matches, improving recall on complex multi-hop queries where either path alone would fail.
-
Bounded Multi-Hop Expansion: Add a configurable expansion mechanism that performs event-to-entity-to-event composition for up to L rounds (default L=1), using reverse SQL joins to discover new entities and events from seed results. This allows the system to traverse reasoning chains while keeping computational cost bounded, scaling to large corpora without performance collapse.
-
LLM-Based Contextual Final Selection: Replace pointwise reranking with an LLM that reads the top-K candidate events and jointly selects up to 5 events that collectively entail the answer. This improves retrieval precision by considering inter-document evidence compatibility, rather than scoring each document independently.
-
Embedding-Model Robustness: Design the retrieval architecture to rely primarily on exact entity joins for structured recall, with embeddings only for semantic fallback. This reduces sensitivity to embedding model quality, allowing the system to maintain high performance even with weaker or smaller embedding models.
-
Graceful Degradation Under Corpus Growth: Use an append-only index that does not require recomputation when new documents arrive. This ensures retrieval quality degrades slowly as the corpus expands, unlike graph-based methods that suffer from fragmentation and stale structures.
What the Improved AI System Can Do:
-
Answer multi-hop questions (e.g.,
Who directed the film starring the actor that won the Oscar in 2019?
) by dynamically linking events across documents at query time, without needing a pre-built knowledge graph. -
Maintain high retrieval accuracy (e.g., 80%+ Recall@5 on MuSiQue) even when the underlying embedding model is swapped for a weaker one, losing only 1.35 points instead of 9.42.
-
Scale to large document collections (e.g., 11,000+ chunks) while retaining 99% of full-pool retrieval quality using only a 500-candidate activation window, enabling efficient deployment in production.
-
Support incremental updates—new documents can be added to the index instantly without recomputing existing records, making it suitable for continuously growing corpora like news feeds or agent memory.
-
Provide explainable evidence chains, as every retrieved piece of evidence remains the original chunk, allowing users to trace the multi-hop reasoning path through shared entities.
Sources
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Message Passing for Hyper-Relational Knowledge Graphs
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
- ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
- Unsupervised Dense Information Retrieval with Contrastive Learning
- StructGPT: A General Framework for Large Language Model to Reason over Structured Data
- ATOM: AdapTive and OptiMized dynamic temporal knowledge graph construction using LLMs
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
- HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
- Generative Representational Instruction Tuning
- Qwen3 Technical Report
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering