SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

arXiv:2608.12129 · cs.CL · Submitted 2026-08-24 · Read on arXiv

Yuchao Wu, Junqin Li, XingCheng Liang, Yongjie Chen, Yinghao Liang, Linyuan Mo, Guanxian Li

Zleap AI

cs.CL

Submitted: 2026-08-24

Updated: 2026-08-25

Code: https://github.com/Zleap-AI/SAG-Benchmark

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges" 1.

Terminology

Summary

Summary of the Paper SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

1. Core Problem and Motivation

The paper addresses the limitations of existing retrieval-augmented generation (RAG) systems for multi-hop question answering. It states that mainstream dense-retrieval implementations remain inherently limited in handling structured constraints and multi-hop reasoning. While graph-based methods like GraphRAG address this by constructing knowledge graphs offline, they often fragment semantics, incur high maintenance, and complicate incremental updates. The central research question posed is: can the structure needed by a query be activated from an append-only index during retrieval, without constructing a global knowledge graph?

2. Proposed Method: SAG

The authors propose SAG (SQL-Retrieval Augmented Generation), a structured retrieval architecture that organizes documents into an event-entity index without building a global knowledge graph.

  • Event-Entity Index: Each chunk is processed independently into one semantically complete event and a set of entities. This pair defines a latent hyperedge that preserves n-ary relations without decomposing them into triples. The index is append-only, allowing new documents to be added without recomputing existing records.

  • Query-Time Dynamic Hyperedges: At query time, SAG treats shared entities as join keys to connect related chunks using SQL joins. This dynamically yields a query-scoped neighborhood of events, and every piece of evidence remains the original chunk throughout.

  • Architecture: The system has an offline indexing phase and an online phase with seed retrieval, query-time expansion, and final selection. It uses SQL for exact entity joins, vector retrieval for semantic similarity, and an LLM for event-entity extraction, query-entity identification, and final selection.

3. Key Technical Details

  • Seed Retrieval: SAG uses two parallel paths: Path A (entity-guided structured recall) identifies query entities and uses SQL joins to find events sharing those entities; Path B (direct event recall) retrieves events by embedding similarity to the query.

  • Expansion: Starting from seed events, SAG uses reverse SQL joins to find new entities and then new events, performing event-to-entity-to-event composition for up to L rounds (default L=1).

  • Final Selection: An LLM reads the top-K candidates and selects up to 5 events that jointly entail the answer. This is a contextual selection step rather than a pointwise reranking. The semantic path fills the remaining slots to return 10 chunks total.

4. Experimental Setup

  • Datasets: MuSiQue (up to 4 hops), 2WikiMultiHopQA, and HotpotQA (2 hops). For corpus growth, they also use NQ.

  • Baselines: Compared against simple retrievers (BM25, Contriever), large embedding models (BGE-Large, GTE-Qwen2, GritLM, NV-Embed), and structure-augmented methods (GraphRAG, LightRAG, HippoRAG 2, HyperGraphRAG, HyperRAG).

  • Unified Configuration: All structure-augmented methods use BGE-Large-EN-v1.5 as the retriever and Qwen3.6-Flash as the reader to isolate architecture from model choice.

5. Main Results

SAG achieves the best retrieval and end-to-end QA performance on all three benchmarks, with gains widening as reasoning-chain complexity increases.

  • Retrieval (Recall@5):

  • On MuSiQue, SAG reaches 80.36%, outperforming the strongest baseline (HippoRAG 2 at 65.13%) by 15.23 points.

  • On 2WikiMultiHopQA, SAG reaches 93.34%, surpassing HippoRAG 2 (90.35%).

  • On HotpotQA, SAG reaches 96.50%, leading HippoRAG 2 (94.35%) by 2.15 points.

  • QA (F1):

  • On MuSiQue, SAG achieves 61.15 F1, exceeding GraphRAG by 7.01 points.

  • On 2WikiMultiHopQA and HotpotQA, SAG leads the strongest baseline by 1.17 and 1.66 points, respectively.

6. Analysis and Ablations

  • Ablation on MuSiQue: Final selection has the largest effect (replacing the LLM with a reranker reduces Recall@5 by 13.25 points). Disabling expansion reduces Recall@5 by 10.95 points. Replacing hyperedge indexing with triple indexing causes a smaller decrease of 2.75 points.

  • Embedding Robustness: SAG is more robust to embedding model changes. Replacing NV-Embed-v2 with BGE-v1.5 reduces HippoRAG 2 Recall@5 by 9.42 points, but SAG only drops by 1.35 points.

  • Scaling: With a candidate pool of 500 (out of 11,656), SAG retains 99% of the full-pool result (79.57% vs 80.36% Recall@5), showing bounded activation limits downstream workloads.

  • Corpus Growth: SAG degrades more slowly than HippoRAG 2 as the corpus grows. On MuSiQue, SAG drops from 87.90% to 82.57% (loss of 5.33 points), while HippoRAG 2 falls from 69.23% to 60.27% (loss of 8.96 points).

7. Limitations

The paper identifies two main limitations:

  1. No Alias Resolution: SAG normalizes and deduplicates entity strings but does not resolve aliases, so it cannot recognize that Apple Inc. and Apple refer to the same entity.

  2. No Temporal Updates: The index is append-only, providing no mechanism for revising or retiring stale events, which is insufficient for agent memory that must represent changing facts and preferences.

8. Conclusion

The paper concludes that SAG turns indexed data into structured context for multi-hop retrieval, achieving the strongest average retrieval and QA performance under unified settings. The authors state that event representation, query-time expansion, and LLM-based final selection each contribute to retrieval quality, and that SAG varies less across embedding models and degrades more slowly as the corpus grows. Future work will explore alias resolution and versioned updates.

Improvements for AI systems

Improvements to AI Systems Based on SAG:

  1. Dynamic Query-Time Structuring: Implement an append-only event-entity index where each document chunk is stored as a semantically complete event with linked entities. At query time, use SQL-style joins on shared entities to dynamically construct query-scoped hyperedges, eliminating the need for pre-built global knowledge graphs. This enables the system to handle multi-hop reasoning without costly offline graph construction or maintenance.

  2. Hybrid Retrieval with Dual Pathways: Integrate two parallel retrieval paths—entity-guided structured recall (using exact SQL joins on query entities) and semantic embedding-based event recall. Combine results to capture both precise relational constraints and fuzzy semantic matches, improving recall on complex multi-hop queries where either path alone would fail.

  3. Bounded Multi-Hop Expansion: Add a configurable expansion mechanism that performs event-to-entity-to-event composition for up to L rounds (default L=1), using reverse SQL joins to discover new entities and events from seed results. This allows the system to traverse reasoning chains while keeping computational cost bounded, scaling to large corpora without performance collapse.

  4. LLM-Based Contextual Final Selection: Replace pointwise reranking with an LLM that reads the top-K candidate events and jointly selects up to 5 events that collectively entail the answer. This improves retrieval precision by considering inter-document evidence compatibility, rather than scoring each document independently.

  5. Embedding-Model Robustness: Design the retrieval architecture to rely primarily on exact entity joins for structured recall, with embeddings only for semantic fallback. This reduces sensitivity to embedding model quality, allowing the system to maintain high performance even with weaker or smaller embedding models.

  6. Graceful Degradation Under Corpus Growth: Use an append-only index that does not require recomputation when new documents arrive. This ensures retrieval quality degrades slowly as the corpus expands, unlike graph-based methods that suffer from fragmentation and stale structures.

What the Improved AI System Can Do:

  • Answer multi-hop questions (e.g., Who directed the film starring the actor that won the Oscar in 2019?) by dynamically linking events across documents at query time, without needing a pre-built knowledge graph.

  • Maintain high retrieval accuracy (e.g., 80%+ Recall@5 on MuSiQue) even when the underlying embedding model is swapped for a weaker one, losing only 1.35 points instead of 9.42.

  • Scale to large document collections (e.g., 11,000+ chunks) while retaining 99% of full-pool retrieval quality using only a 500-candidate activation window, enabling efficient deployment in production.

  • Support incremental updates—new documents can be added to the index instantly without recomputing existing records, making it suitable for continuously growing corpora like news feeds or agent memory.

  • Provide explainable evidence chains, as every retrieved piece of evidence remains the original chunk, allowing users to trace the multi-hop reasoning path through shared entities.

Sources

Related papers