Knowledge Pull Requests for Continual Document Authoring
cs.CL
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Code: https://github.com/alexmartin1722/kpr
Code: https://github.com/alexmartin1722/kpr
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable.
Terminology
Abstract
We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content, producing a ChangeLog that separates what knowledge changes (claim proposal) from how the text changes (document diff). We evaluate KPRs on revising Wikipedia across languages and updating query-driven reports on RAGTIME. KPRs integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated. A KPR-revised article also grounds question answering better than a frontier model with search, which does not surface knowledge documented only in other languages.
Sources
- The Llama 3 Herd of Models
- MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources
- MegaWika: Millions of reports and their sources across 50 diverse languages
- LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
- AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
- Mixtral of Experts
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Overview of the TREC 2025 RAGTIME Track
- Olmo 3
- Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
- Auto-ARGUE: LLM-Based Report Generation Evaluation
- Gemma 4 Technical Report
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering