MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models
cs.CL, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
- ThinkEval: Practical Evaluation of Knowledge Leakage in LLM Editing using Thought-based Knowledge Graphs
- Improving language models by retrieving from trillions of tokens
- EvoEdit: Lifelong Free-Text Knowledge Editing through Latent Perturbation Augmentation and Knowledge-driven Parameter Fusion
- UniEdit: A Unified Knowledge Editing Benchmark for Large Language Models
- Beyond Memorization: A Rigorous Evaluation Framework for Medical Knowledge Editing
- Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
- AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models
- Gemma 3 Technical Report
- Model Editing at Scale leads to Gradual and Catastrophic Forgetting
- REALM: Retrieval-Augmented Language Model Pre-Training
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Patching open-vocabulary models by interpolating weights
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- Dense Passage Retrieval for Open-Domain Question Answering
- RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge
- Locating and Editing Factual Associations in GPT
- Memory-Based Model Editing at Scale
- WARM: On the Benefits of Weight Averaged Reward Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering