Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/jmendi/eolas_irradiated_materials
License: http://creativecommons.org/licenses/by/4.0/
The gist: The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels.
Terminology
Abstract
The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels. However, essential data necessary for these models and simulations are often embedded in scientific literature as unstructured text, limiting reusability and posing challenges for researchers seeking to leverage existing knowledge effectively. While extracting structured data from unstructured text using large language models is gaining popularity, traditional methods typically generate key-value pairs data with straightforward schemas. In contrast, we introduce eolas, a modular pipeline that uses large language models to automatically transform scientific documents into knowledge graphs aligned with a specified ontology. We demonstrate eolas effectiveness in extracting useful information for scientists studying materials designed to endure the extreme temperatures and radiation levels found in fusion reactors. While a human expert might spend between thirty to ninety minutes extracting relevant data from an article, eolas can generate high-quality knowledge graphs in just a few minutes. These are presented in a tabular format with faceted navigation for easy human validation. Additionally, we introduce the first benchmark dataset designed to assess large language models capabilities in constructing knowledge graphs within the domain of irradiated materials. The analysis of 168 experiments using our dataset, various large language models and prompting techniques provides key insights that we summarize into practical guidelines for effectively extracting knowledge graphs aligned with an input ontology.
Sources
- Automated, LLM enabled extraction of synthesis details for reticular materials from scientific literature
- iText2KG: Incremental Knowledge Graphs Construction Using Large Language Models
- Exploring In-Context Learning Capabilities of Foundation Models for Generating Knowledge Graphs from Text
- gpt-oss-120b & gpt-oss-20b Model Card
- Why Language Models Hallucinate
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Graph Retrieval-Augmented Generation: A Survey
- Docling Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection