Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
cs.LG, cs.AI, cs.CL, cs.IR
Submitted: 2025-11-06
Updated: 2026-09-21
Comments: Published in Transactions on Machine Learning Research (TMLR)
Journal ref: Transactions on Machine Learning Research (TMLR), 09/2026
Code: https://github.com/graphcore-research/synth-kgqahttps:
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs.
Terminology
Abstract
Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various solutions have been proposed, a comparison of methods is difficult due to the lack of challenging QA datasets with ground-truth targets for graph retrieval. We present SynthKGQA, an LLM-powered framework for generating high-quality Knowledge Graph Question Answering datasets from any Knowledge Graph, providing the full set of ground-truth facts in the KG to reason over questions. We show how, in addition to enabling more informative benchmarking of KG retrievers, the data produced with SynthKGQA also allows us to train better models.We apply SynthKGQA to Wikidata to generate GTSQA, a new dataset designed to test zero-shot generalization abilities of KG retrievers with respect to unseen graph structures and relation types, and benchmark popular solutions for KG-augmented LLMs on it.
Sources
- The Llama 3 Herd of Models
- Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering Datasets
- zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
- Recent Advances in Named Entity Recognition: A Comprehensive Survey and Comparative Study
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- GPT-4 Technical Report
- Graph Retrieval-Augmented Generation: A Survey
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks