EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models
cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted in EMNLP Industry track 2026
Journal ref: EMNLP Industry track 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large
Terminology
Abstract
For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question-answering on SLMs often operates under inevitable constraints: small literature collections, fragmented evidence, limited context window and reasoning abilities. We propose the Evidence-Grounded Typed Knowledge Graph (EGT-KG), a retrieval framework to improve information retrieval with local SLMs. We assessed three question-answering settings: a vanilla Retrieval-Augmented Generation (RAG) workflow and two EGT-KG workflows: an automatically generated relation schema (AS) and an expert-defined relation schema (ES). Our experiments were evaluated with a six-dimensional evaluation framework (S3CRF: Soundness, Correctness, Completeness, Conciseness, Relevance, Fluency) on a Biopolymer-bound Soil Composite literature benchmark, showing that EGT-KG outperforms the vanilla RAG method in most settings, with the best improvement from llama3:8b: a Final Score of 70.37 (+14.67%) and 68.82 (+12.14%) by AS/ES EGT-KG variants.
Sources
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- AutoKG: Efficient Automated Knowledge Graph Generation for Language Models
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
- Automated Construction of Theme-specific Knowledge Graphs
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions
- Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Prompt Injection attack against LLM-integrated Applications
- Small Language Models: Survey, Measurements, and Insights
- A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation
- LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale Corpora
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection