Building Knowledge Graphs Towards a Global Food Systems Datahub

arXiv:2502.19507 · cs.AI · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Building Knowledge Graphs Towards a Global Food Systems Datahub".

Jane: The paper was written by Nirmal Gelal, Aastha Gautam, Sanaz Saki Norouzi, Nico Giordano, Claudio Dias da Silva Jr et al. from Kansas State University and University of Missouri.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone! Today we're diving into a paper that's got a big, ambitious title: "Building Knowledge Graphs Towards a Global Food Systems Datahub." Jane, when you first saw that title, what went through your head?

Jane: Honestly, Tom, my first thought was, "That's a mouthful." But once I started reading, it made so much sense. We're talking about a global system for food, and they want to build a "datahub" — basically a giant, organized library of information about how we grow food sustainably.

Tom: And they're using knowledge graphs to do it. For our listeners who might not be familiar, can you break that down?

Jane: Sure. Imagine a giant web of connected facts. Instead of storing data in separate spreadsheets that never talk to each other, a knowledge graph links them together. So you can ask questions like, "What's the relationship between nitrogen levels in the soil and wheat disease resistance?" and the graph can show you the path between those two concepts.

Tom: And that's exactly what this team from Kansas State University is doing. They're starting with wheat — a huge deal for global food security — and they're building this graph to capture everything from nitrogen management to disease control.

Jane: Right. And the "global" part is key. They want this to be a model that can eventually apply to other crops, other regions, other farming systems. It's not just about Kansas wheat; it's about creating a framework that the whole world can use.

Tom: Lu, you're our AI researcher. What excites you about this title and the vision behind it?

Lu: The ambition, Tom. A "global food systems datahub" means we're finally treating agricultural data as a first-class citizen in the AI world. Right now, so much of this data is siloed — weather data here, soil data there, disease reports somewhere else. This paper is about breaking those silos down.

Jane: And that's where the knowledge graph really shines. It's not just a database; it's a way to represent relationships. You can ask it not just "what happened?" but "why did it happen?" and "what might happen next?"

Tom: Meng, you're the engineer. What's your gut reaction to a project this size?

Meng: My first question is always, "How do you keep it from falling apart?" Building a knowledge graph is one thing; building one that's scalable, maintainable, and actually useful for farmers and researchers is another. But the modular design they mention in the abstract gives me some hope — they're not trying to boil the ocean all at once.

Tom: So they're starting small with wheat, but thinking big. That's a smart approach. What's the next piece of this paper that we should dig into?

Jane: I think we should talk about the methodology — how they're actually building this thing. It's called KNARM, and it's all about getting humans and machines to work together.

Summary: Tom: So we're back, and we're still talking about "Building Knowledge Graphs Towards a Global Food Systems Datahub." Jane, you mentioned KNARM — that's the methodology they're using. What's the big idea there?

Jane: KNARM stands for Knowledge Acquisition and Representation Methodology. The key insight is that you can't just scrape data and throw it into a graph. You need domain experts — agronomists, farmers, pathologists — to tell you what matters. The methodology is a structured way to interview those experts, extract their knowledge, and turn it into a formal schema.

Tom: So it's like building a blueprint for the knowledge graph, and the blueprint comes from talking to people who actually know wheat.

Jane: Exactly. And they do it in stages. First, they analyze existing data sources. Then they do unstructured interviews with experts — just open conversations. Then they look at existing ontologies, which are like pre-built vocabularies for specific domains, to see what they can reuse.

Lu: That's the part I find really smart, Tom. They're not reinventing the wheel. There are already ontologies for weather, for soil, for crop diseases. The paper mentions Agroportal, which is a hub for agricultural vocabularies. By reusing those, they ensure their graph can talk to other systems.

Meng: But reusing ontologies can get messy. You have different naming conventions, different levels of detail. How do you make sure the pieces fit together?

Jane: That's where the structured interviews come in. They ask experts very specific, close-ended questions to validate the schema. And they develop "competency questions" — if the graph can answer these questions, the schema is complete.

Tom: Can you give us an example of a competency question?

Jane: Sure. Something like, "What nitrogen application rate is recommended for a wheat crop with a yield expectation of eighty bushels per acre under drought stress?" If the knowledge graph can answer that, then the schema is capturing the right relationships.

Meng: That's a practical test. I like that. It's not just about academic completeness; it's about whether the system actually works for decision-making.

Lu: And the paper shows they're already applying this to two key areas: nitrogen management and disease management. Those are two of the biggest levers for sustainable wheat production.

Tom: So the summary is: they've got a methodology, they've got a schema in progress, and they're starting with nitrogen and disease. What's the coolest part of the actual schema they've built so far?

Jane: I think the disease management module is fascinating. They've broken it down into fungi, viruses, and bacteria, and then for fungi, they have three sub-modules: chemical strategies, cultural practices, and host genetics. It's a really clean way to organize a complex problem.

Tom: And that's where we're headed next — the actual improvements and innovations this paper brings to the table.

Improvements: Tom: We're back with "Building Knowledge Graphs Towards a Global Food Systems Datahub," and now I want to get into the improvements this paper suggests. Jane, what's the biggest leap forward here?

Jane: For me, it's the sustainability module. They've taken two established frameworks — SMART and IDEA — and mapped them into their knowledge graph. These frameworks break sustainability down into dimensions, themes, and sub-themes. By integrating those into the graph, they can now ask questions like, "How does nitrogen management on a specific farm impact its social sustainability score?"

Tom: So they're not just tracking yield and disease — they're tracking the whole picture of what "sustainable" means.

Jane: Exactly. And they go one step further. They map those sustainability themes to the United Nations Sustainable Development Goals. So you can trace a specific farming practice all the way up to a global goal like "Zero Hunger" or "Climate Action."

Lu: That's the kind of cross-domain thinking that makes this paper exciting. It's not just an agricultural ontology; it's a bridge between agriculture and global policy. That's rare.

Meng: But I want to know about the practical side. They mention using GraphDB and Python scripts to populate the data. How does that actually work in practice?

Jane: They've set up a server at Kansas State, and they're using GraphDB as their database. It supports RDF and SPARQL, which are the standard languages for knowledge graphs. They're writing scripts to take flat files — like weather data or soil data — and automatically convert them into graph format.

Meng: And that's where the scalability question comes in. If you're ingesting data from thousands of farms, you need that process to be robust.

Jane: Right. And that's why they're being careful about data validation. They have domain experts reviewing the schema, and they're using reasoners — software that checks for logical consistency — to catch errors.

Tom: What about the human side? They mentioned a Sustainability Workshop they organized. What was that about?

Jane: That's one of my favorite parts. They brought in farmers, bakers, millers — the whole wheat value chain — and asked them what sustainability means to them. That's invaluable data that you can't get from a satellite or a soil sensor.

Lu: And that's what makes this different from other knowledge graph projects. It's not just top-down; it's bottom-up. The people who actually grow and use wheat are shaping the ontology.

Meng: So the improvement here is really about integration — integrating diverse data sources, integrating expert knowledge, and integrating sustainability frameworks into one coherent system.

Tom: And that integration is what could make this a game-changer. But what are the challenges? What's standing in the way?

Conclusion: Tom: We're wrapping up our discussion on "Building Knowledge Graphs Towards a Global Food Systems Datahub," and I want to get final thoughts from everyone. Jane, what's the takeaway for our listeners?

Jane: The takeaway is that we're finally getting serious about structuring agricultural data. This paper lays out a clear path for building a knowledge graph that can handle the complexity of sustainable food production — not just for wheat, but eventually for all crops.

Tom: Lu, what excites you most about where this could go?

Lu: The potential for AI and machine learning is enormous. Once you have a well-structured knowledge graph, you can train models to predict yields, detect disease outbreaks early, and even recommend optimal planting times. The graph becomes the foundation for a whole new generation of agricultural intelligence.

Meng: And from an engineering standpoint, I appreciate that they're using established standards — RDF, OWL, SPARQL. That means this isn't a one-off project; it's something that can be built upon and integrated with other systems.

Tom: Lalam, you're our in-house language model. What's your perspective on the cultural impact of this work?

Lalam: I think the cultural impact is about democratizing knowledge. Right now, a farmer in Kansas and a researcher in Kenya might have completely different information about wheat production. A global datahub could level that playing field, giving everyone access to the same insights and best practices.

Jane: And that's the real promise here. It's not just about technology; it's about creating a shared resource that can help feed the world more sustainably.

Tom: Well said. We've covered the title, the methodology, the improvements, and the big picture. This paper is a solid step toward a future where agricultural data is open, connected, and actionable.

Jane: And we're excited to see where this research goes. Thanks for joining us, everyone. We'll be back soon with another paper to break down.

Tom: Until next time, keep asking questions and keep learning. Goodbye, everyone!

Nirmal Gelal, Aastha Gautam, Sanaz Saki Norouzi, Nico Giordano, Claudio Dias da Silva Jr, Jean Ribert Francois, Kelsey Andersen Onofre, Katherine Nelson, Stacy Hutchinson, Xiaomao Lin, Stephen Welch, Romulo Lollato, Pascal Hitzler, Hande Küçük McGinty

Kansas State University · University of Missouri

cs.AI

Submitted: 2026-08-14

Updated: 2026-08-17

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 41/100

Key concepts

Knowledge Graph
A knowledge graph is described as a giant web of connected facts. Instead of separate spreadsheets, it links data points together, allowing users to ask complex questions about relationships between concepts, such as soil nitrogen levels and wheat disease resistance.
KNARM
This stands for Knowledge Acquisition and Representation Methodology. It is the structured process used to build the graph by interviewing domain experts (like agronomists) to extract their knowledge and formalize it into a schema, rather than just scraping raw data.
Global Food Systems Datahub
The goal of the project is to create a worldwide, organized repository of information on how food is grown sustainably. It aims to be a universal framework applicable beyond specific regions or crops, helping to break down siloed agricultural data.
Ontology
An ontology acts like a pre-built vocabulary for a specific domain (like weather or soil). By reusing existing ontologies, the knowledge graph can ensure its data can communicate and connect with other systems and vocabularies.

Terminology

Summary

Summary

This paper addresses the lack of a standardized framework for storing and analyzing data on sustainable agricultural practices, specifically within wheat production, which is a key pillar of global food security. The authors note that while the Food and Agriculture Association of the United Nations reports that global food demand will increase by 50 percent by 2050, there is currently no structured method for representing sustainability information in wheat production. To address this gap, the authors are building a set of ontologies and Knowledge Graphs (KGs) that encode knowledge associated with sustainable wheat production using formal logic. The data for these knowledge graphs are collected from public data sources, experimental results collected at Kansas State University, and a Sustainability Workshop organized by the authors, which helped collect input from stakeholders throughout the wheat value chain. The modeling of the ontology (i.e., the schema) for the Knowledge Graph is in progress with the help of domain experts, following a modular structure using the KNARM methodology. The paper presents preliminary results and schemas of the Knowledge Graph and ontologies.

The authors explain that their objective is to build a unified KG (combination of a set of KGs) on a modular framework, where each module focuses on a distinct aspect of sustainable wheat production, making data addition and removal manageable and enhancing scalability. This modular structure enables integration of various components of the wheat value chain and allows seamless integration with other agricultural datasets. As part of preliminary results, the primary objective has been to encode knowledge on nitrogen management and disease management in wheat, providing researchers, agronomists, and policymakers with data-driven insights to support sustainable practices.

In the literature review, the authors discuss that knowledge graphs have gained popularity due to their benefits including scalability and structuring heterogeneous data into a single structure. They mention several existing efforts in agriculture, including Agroportal, which serves as a hub containing vocabularies and ontologies related to agriculture. They also discuss the Ploutos data-sharing architecture, which is based on three principles: reuse of existing semantic standards, integration with legacy systems, and a distributed architecture where stakeholders control access to their own data. The authors also review domain-specific ontologies such as the Crop Disease Ontology, the Environment Ontology (ENVO), BIMERR Weather Ontology, Weather Ontology, Ontology for Meteorological Sensors, and the Phenotype Quality Ontology (PATO). They also mention KnowWhereGraph, which they describe as the currently largest public geo-knowledge graph, covering agriculture-relevant data including soil health data, land use, and land cover data. The authors note that these existing ontologies face limitations in integration with knowledge from other domains, as they are primarily designed to address specific problems within well-defined domain boundaries.

The methodology section describes the use of the Knowledge Acquisition and Representation Methodology (KNARM), which addresses the knowledge acquisition bottleneck by incorporating human expertise. KNARM focuses on building an ontology or knowledge graph with the end applications and available datasets in mind, aiming to build templates based on concrete use cases. The methodology includes nine steps: Sublanguage Analysis, Unstructured Interview, Sub-language Recycling, Meta-Data Creation and Knowledge Modeling, Structured Interview, Knowledge Acquisition Validation, Database Formation, Semi-Automated Ontology Building, and Ontology Validation and Evaluation. In the Sublanguage Analysis step, the authors examine existing data sources, analyze patterns within the data, and define relationships between existing knowledge sources. The Unstructured Interview step involves informal interviews with domain experts. In Sub-language Recycling, existing ontologies are collected and evaluated to facilitate reuse of vocabularies. Meta-Data Creation and Knowledge Modeling involves listing abstract concepts related to the data hub. The Structured Interview step involves close-ended questions for domain experts and formulation of competency questions to ensure schema completeness. Knowledge Acquisition Validation is an initial feedback loop where outputs are presented to domain experts to identify gaps. Database Formation involves choosing a database type, and the authors opted for Graph-DB due to structural compatibility with the schema, scalability, and inherent graph-based architecture. Semi-Automated Ontology Building utilizes tools such as RDFLib in Python, Jena in Java, and ROBOT command-line tool. The final step, Ontology Validation and Evaluation, involves testing whether the ontology accurately represents intended information, including answering competency questions and performing information retrieval tasks using SPARQL.

The ontology and schema design process is ongoing, with an initial version addressing nitrogen and disease management within the wheat domain. The nitrogen management module connects nitrogen demand and supply within a nitrogen management strategy and examines the relationship between crop yield expectations and nitrogen requirements. The authors note that crop yield expectation correlates directly with nitrogen demand, as higher yields necessitate greater nitrogen inputs. Key factors supporting maximum crop yields, such as Flowering Date, Freeze Damage, and Heat Stress, are illustrated in the schema. The disease management module focuses on reducing the impact of pathogens that can significantly lower wheat yields. Common wheat pathogens—fungi, viruses, and bacteria—are depicted in the top-level schema, with current focus on fungi, detailing their types and potential treatment strategies including Chemical Strategy, Cultural Practices, and Host Genetics. Each of these treatment strategies functions as a self-contained sub-module, encapsulating concepts and relationships within its specific sub-domain.

Schema design requires multiple iterations, and the authors involve domain experts throughout the process, conducting both structured and unstructured interviews to iteratively build the schema until it adequately captures the desired information. Before designing the schema, the authors developed a list of competency questions, which define the scope and goals of the ontology. They are using Protégé, an open-source ontology development tool, to build the ontology based on the generated schema. The Web Ontology Language (OWL) will be employed to axiomatize knowledge within the domain of agriculture, with fundamental axioms including definitions of domains, ranges, constraints, and equivalences formally encoded. The authors plan to infer new knowledge from the knowledge base using reasoners such as HermiT, Pellet, and other existing reasoners.

For data collection and preparation, the authors are examining various pre-existing ontologies within the wheat domain or related domains, including the Wheat Trait Ontology (WTO), Environment Ontology (ENVO), Crop Dietary Nutrition, and KnowWhereGraph. Reusing established ontologies offers advantages such as saving time, ensuring consistency, and enhancing interoperability. The authors have identified several key data sources, including weather, soil, and landscape datasets, which directly impact each stage of wheat production, influencing factors such as quality, yield, and sustainability. They organized a Sustainability Workshop inviting key stakeholders from the wheat value chain, including farmers, bakers, millers, and other domain experts, and the data gathered from these expert contributions is currently under analysis. They also plan to incorporate insights from the workshop into the knowledge graph and are integrating ground-truth experimental data collected by professors at Kansas State University.

For database and data population, the authors have set up a dedicated server at Kansas State University to store all knowledge graph data in GraphDB, a specialized database and semantic repository designed for storage, management, and querying of knowledge graphs. They are developing Python scripts utilizing libraries that work with RDF data formats to populate data from flat files efficiently into GraphDB.

The authors also developed a sustainability module by mapping two established sustainability frameworks, SMART and IDEA. These frameworks provide a structured hierarchy for sustainability metrics, organizing them into dimensions, themes, and sub-themes. Using these references, the authors systematically aligned their sustainability metrics with the corresponding dimensions, themes, and sub-themes. To implement this module, they utilized ROBOT, a flexible command-line tool for working with ontologies in OWL. A domain-user-friendly template file (CSV) was created, which ROBOT efficiently converts into an OWL format. By integrating these sustainability dimensions, themes, and sub-themes with the United Nations Sustainable Development Goals, the expanded graph establishes mappings between key sustainability concepts and processes associated with wheat cultivation.

In the conclusion, the authors state that their DataHub provides a robust foundation for addressing numerous downstream Artificial Intelligence (AI) and Machine Learning (ML) tasks. It incorporates a comprehensive dataset encompassing weather patterns, soil characteristics, and other key agricultural variables, enabling stakeholders to optimize decision-making processes such as planting and harvesting schedules, implement preventive measures, and minimize losses. Examples of ML models well-suited to leverage the DataHub include regression models for predicting yields based on parameters such as nitrogen levels, weather patterns, and soil conditions; classification models for identifying disease outbreaks; clustering algorithms for segmenting fields into zones to facilitate targeted interventions; and reinforcement learning models to dynamically optimize irrigation and fertilizer applications.

The authors acknowledge limitations of knowledge graphs, including time complexity during reasoning and inference processes in large graphs, and knowledge validation challenges. They also note the challenge of semantic and naming inconsistency when creating a datahub combining a set of KGs, and plan to use manual checking and mapping rules to ensure proper ontology alignment, including axioms like equivalence, subclass, and instance mapping. They also plan to use reasoners to check logical inconsistencies in the merged ontology. Future work aims to transform the project into a fully operational datahub of wheat that aids in the global food system datahub, with planned expansions including extending the disease management module to cover fungi and bacteria, and incorporating additional modules focusing on weather, soil, drought, and other related concepts. The goal is to develop a comprehensive DataHub that encapsulates all fundamental concepts within the wheat value chain, spanning from farm to table. The methodologies employed are designed for adaptability, allowing similar structures to be applied to other crops.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:


Improvement: I will build a modular, ontology-driven knowledge graph (KG) using the KNARM methodology, with initial modules for nitrogen management and disease management (fungi, with sub-modules for chemical, cultural, and host genetics strategies). The schema will be encoded in OWL using Protégé, with axioms for domains, ranges, constraints, and equivalences, and validated via competency questions and SPARQL queries.

What the improved AI system can do:

  • Answer complex queries such as: Which cultural practices reduce fungal disease incidence in wheat under high nitrogen conditions? or What is the optimal nitrogen application rate for a given yield expectation and soil type?

  • Integrate heterogeneous data (weather, soil, landscape, experimental field data) into a unified, queryable structure.

  • Perform logical reasoning (using HermiT or Pellet) to infer new knowledge, e.g., detecting inconsistencies in management recommendations or deriving implicit relationships between disease management strategies and sustainability outcomes.

Improvement: I will implement a sustainability module that maps the SMART and IDEA frameworks (dimensions, themes, sub-themes) to the United Nations Sustainable Development Goals (SDGs), using ROBOT to convert CSV templates into OWL axioms.

Improvement: I will develop a semi-automated data population pipeline using Python (RDFLib) and GraphDB, with scripts that convert flat files (e.g., weather, soil, experimental data) into RDF triples, following the ontology schema. I will also incorporate data from the Sustainability Workshop (stakeholder input) and Kansas State University field experiments.

Improvement: Using the populated KG, I will train and deploy machine learning models that leverage the structured data for downstream tasks, including:

  • Regression models for yield prediction based on nitrogen levels, weather patterns, soil properties, and disease incidence.

  • Classification models for early detection of fungal disease outbreaks using environmental and management features.

  • Clustering algorithms to segment fields into management zones for targeted interventions.

  • Reinforcement learning for dynamic optimization of irrigation and fertilizer application.

Improvement: I will implement ontology alignment techniques (manual mapping rules and automated equivalence/subclass axioms) to merge the wheat KG with existing ontologies (ENVO, WTO, PATO, KnowWhereGraph) and resolve naming inconsistencies. I will use reasoners to check logical consistency post-merger.

Improvement: I will build a natural language interface (using the KG’s SPARQL endpoint) that translates user questions into formal queries and returns results with explanations, addressing the limitation that non-technical users struggle with graph query languages.

Improvement: I will design the ontology and KG architecture to be modular and crop-agnostic, with the wheat-specific modules (nitrogen, disease) easily replaceable or extendable for other crops (e.g., corn, soybean).

These improvements directly address the paper’s stated gaps (lack of standardized vocabularies, modular KG design, and integration of expert knowledge) and its future work (expanding to other crops, enabling AI/ML tasks). The resulting AI system will be a robust, queryable, and reasoning-capable datahub that supports data-driven decision-making for sustainable wheat production and beyond.

Abstract

Sustainable agricultural production aligns with several sustainability goals established by the United Nations (UN). However, there is a lack of studies that comprehensively examine sustainable agricultural practices across various products and production methods. Such research could provide valuable insights into the diverse factors influencing the sustainability of specific crops and produce while also identifying practices and conditions that are universally applicable to all forms of agricultural production. While this research might help us better understand sustainability, the community would still need a consistent set of vocabularies. These consistent vocabularies, which represent the underlying datasets, can then be stored in a global food systems datahub. The standardized vocabularies might help encode important information for further statistical analyses and AI/ML approaches in the datasets, resulting in the research targeting sustainable agricultural production. A structured method of representing information in sustainability, especially for wheat production, is currently unavailable. In an attempt to address this gap, we are building a set of ontologies and Knowledge Graphs (KGs) that encode knowledge associated with sustainable wheat production using formal logic. The data for this set of knowledge graphs are collected from public data sources, experimental results collected at our experiments at Kansas State University, and a Sustainability Workshop that we organized earlier in the year, which helped us collect input from different stakeholders throughout the value chain of wheat. The modeling of the ontology (i.e., the schema) for the Knowledge Graph has been in progress with the help of our domain experts, following a modular structure using KNARM methodology. In this paper, we will present our preliminary results and schemas of our Knowledge Graph and ontologies.

Related papers