Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering
cs.CL, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-18
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) suffer from a long-tail deficit: culturally specific facts, particularly those concerning underrepresented regions such as Latin America, appear too rarely in pretraining
Terminology
Abstract
Large language models (LLMs) suffer from a long-tail deficit: culturally specific facts, particularly those concerning underrepresented regions such as Latin America, appear too rarely in pretraining corpora to be reliably memorized. Retrieval-Augmented Generation (RAG) addresses this by grounding generation in external text, but structured alternatives such as Knowledge Graphs (KGs) offer tighter control over what enters the context, along with potential gains in explainability and updatability. We benchmark Graph-RAG against standard RAG on LatamQA, a culturally grounded multiple-choice dataset spanning eight thematic categories. The graphs are built end-to-end from Wikipedia articles with KGGen, a recent open-domain extractor, without manual curation in our main setting. G-Retriever is competitive with RAG and reduces the error of the base LLM by 72% with a standard KG and 78% with a benchmark-aware variant, the gap to RAG narrowing further as the graph is oriented toward task-relevant content. The trained projection transfers zero-shot to Portuguese without target-language fine-tuning, indicating multilingual reach.
Sources
- GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
- Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America
- Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answering
- KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
- Qwen2.5 Technical Report
- ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- Multilingual E5 Text Embeddings: A Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering