Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-28
Comments: This work has been accepted to EMNLP 2026 Findings
Code: https://github.com/LARK-NLP-Lab/Surgical-Alignment
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Biomedical knowledge graphs (KGs) offer structured medical knowledge that can ground large language model (LLM) reasoning in clinical diagnosis application, yet how KG signal should be integrated
Terminology
Abstract
Biomedical knowledge graphs (KGs) offer structured medical knowledge that can ground large language model (LLM) reasoning in clinical diagnosis application, yet how KG signal should be integrated into LLMs remains an open question. We present a systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs. At the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior. We introduce Gradient Intervention Density (GID) and Gradient Distortion (GD) to measure how broadly an optimizer modifies the pretrained model. GID and GD together reveal a clear divide: KG-judgment training under KL regularization produces sparse, localized updates (a regime we term as surgical alignment), while task-specific SFT produces dense ones. A controlled ablation shows that the objective and KL contribute to sparsity independently, and the paradigms that produce sparse updates also improve reasoning quality, even when their in-domain accuracy is lower than task-specific SFT. Assessing KG-LLM integration thus requires complementing accuracy with optimization-geometry diagnostics. Our implementation can be found at https://github.com/LARK-NLP-Lab/Surgical-Alignment.
Sources
- Optuna: A Next-generation Hyperparameter Optimization Framework
- ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
- RM-R1: Reward Modeling as Reasoning
- Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- Reward Reasoning Model
- Reasoning with Graphs: Structuring Implicit Knowledge to Enhance LLMs Reasoning
- LoRA: Low-Rank Adaptation of Large Language Models
- medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs
- MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
- Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
- Process Reward Models That Think
- Medical Hallucinations in Foundation Models and Their Impact on Healthcare
- MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- Qwen2.5 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DDXPlus: A New Dataset For Automatic Medical Diagnosis
- Gemma: Open Models Based on Gemini Research and Technology
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering