LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge
cs.AI, cs.DL, cs.IR
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: 15 (main text) + 6 (SM) pages, 4 + 1 figures
Code: https://github.com/deepseek-ai/deepseek-harness
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence.
Terminology
Abstract
Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process scientific knowledge compilation and implement it in ASKS, the Agent-Driven Scientific Knowledge System. For each source, an LLM produces a readable Wiki view and machine-facing semantics. Deterministic checks convert the latter into a document-local GraphDelta, and embedding geometry together with explicit graph rules integrates the proposed changes into persistent state. Each ingest is an inspectable state transition over accumulated knowledge, with compiled Wiki and graph views linked to the preserved source record. We examine this process by chronologically compiling 56 published papers from one research program. Branch survival, cross-paper support, lineage, coverage, and churn yield a source-traceable author research portrait centered on tensor-network methods, with branches into quantum many-body research, tensor-network machine learning, and quantum-AI-oriented directions. In this run, higher-level Hub organization remains stable and low-churn. Canonical-node growth is predominantly additive. Graph-level measurements and navigation paths retain links to the source records from which they were compiled.
Sources
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- MinerU: An Open-Source Solution for Precise Document Content Extraction
- MiniMax Sparse Attention
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection