K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

arXiv:2605.09635 · cs.CL · Submitted 2026-05-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs".

Tom: Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks only measure factual recall,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, we're tuning into the latest from arXiv today with Hao Liang and his team on their paper titled "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs." The main idea here is addressing a big gap in how we test and train AI for education because current benchmarks only measure factual recall. It claims that to build truly effective educational AI, we need to understand curriculum cognition, which is the structured way knowledge is organized and presented, like knowing prerequisites or how concepts visually connect.

Jane: That makes sense, Tom; it sounds like they're moving beyond just asking a question and checking if the answer is right. They are focusing on whether an AI understands the structure of what it's learning, which is a really important concept for real teaching applications.

Lu: From my perspective at Tsinghua, this knowledge graph approach could unlock entirely new ways to model complex educational pathways; think about how we can map out entire disciplines and see the connections that aren't immediately obvious in a simple Q andA format.

Meng: I wonder if this structural understanding translates to actual practical applications, like creating tutoring systems that can dynamically suggest the next best concept based on what a student has already mastered, rather than just guessing the right answer for one question.

Lalam: If we can teach an AI to understand the "why" and "how" of curriculum structure through this graph, it could fundamentally improve how educational AI models are developed and deployed across different subjects.

Tom: Exactly, Lalam; they're building a resource that does both benchmarking for testing and training data synthesis for teaching these complex relationships to the AI. They claim this K12-KGraph, which they extracted from official Chinese K–twelve textbooks covering math, physics, chemistry, and biology from primary through high school levels, provides the necessary structure <ref:2605.09635#pg0>.

Jane: It’s fascinating that they are extracting this graph by going through a five-stage automatic construction pipeline starting with OCR parsing of the textbooks into structured Markdown and finishing with manual verification by domain experts to ensure correctness. That sounds like a really thorough way to build something reliable.

Lu: The construction pipeline itself is interesting because it handles both the textual curriculum structure—like concept and skill nodes—and the multimodal aspect by capturing visual grounding through nodes like Figure and VisualElement, which lets the AI see how things are presented visually alongside the text.

Meng: From an engineering standpoint, that extraction process sounds intensive; getting those structured edges with evidence citations or confidence scores for every single connection adds a layer of data integrity we don't always see in standard datasets. I want to know how scalable this construction pipeline is when we move it to other curricula or languages.

Paper summary: Lalam: The multimodal component, specifically capturing the visual elements alongside the text, seems particularly powerful because it allows the AI to grasp the connection between a theoretical concept and how it's physically illustrated in a textbook.

Tom: And that’s where they take it further with K12-Bench, which is a twenty-three thousand six hundred forty-question multi-select benchmark designed around five task families like Grounding, Prerequisite Reasoning, Neighbor Recommendation, Evidence Chain, and Locate. This benchmark structure ensures that every sample can be traced back to a specific subgraph because the correct answers are true graph neighbors and the distractors are drawn from structurally proximate but incorrect nodes.

Jane: That systematic control over difficulty and coverage through subgraph analysis is really smart; it moves the evaluation beyond simple factual recall to testing deeper reasoning skills embedded in the curriculum structure.

Lu: The way they design K12-Bench using graph-derived templates, where correct answers are defined as true graph neighbors, gives us a precise mechanism to probe exactly what kind of structural understanding an LLM is actually exhibiting when it solves a problem related to curriculum cognition.

Meng: If we can control the difficulty by manipulating the structure of these subgraphs, that suggests we can train models much more effectively than just feeding them raw text and answers; it’s about teaching them the underlying logic of the knowledge organization itself.

Lalam: This systematic approach to building K12-Bench seems like a perfect bridge between creating a large dataset and creating targeted training signals for educational AI.

Tom: Moving on to K12-Train, they developed this KG-guided supervised fine-tuning corpus with seven thousand three hundred thirty-five samples explicitly designed to teach curriculum cognition through three different paths: node-grounded QA targeting properties like definitions or formulas, edge-grounded QA forcing the model to articulate relations like prerequisites for, and exercise-assessment QA using deterministic templates.

Jane: It’s interesting how they created these different types of questions—one focusing on what the node *is*, another on why one thing must come before another, and a third that uses clear test edges to guarantee factual grounding. That covers a lot of ground for learning structure.

Lu: The partition of this data into K12-Train-Text with two thousand two hundred sixty-seven text-only samples and K12-Train-MM with five thousand sixty-eight multimodal samples shows that the textual and visual supervision are complementary in building this training resource.

Meng: I'm looking at the experimental results now; they mentioned that under a strictly matched two thousand three hundred-sample SFT budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval.

Paper summary: Lalam: That finding suggests that even without the visual component in training, focusing heavily on the explicit structural relationships captured in text is very effective for improving performance on specific educational tasks.

Tom: But they also showed that for Vision-Language Models, K12-Train-Full achieves the best overall performance on Gaokao-MM and MDK12-Bench among all compared training configurations, which points to the multimodal data being quite valuable when it’s available.

Jane: It seems the advantage comes from encoding those explicit relationships—the structural grounding—and having a pedagogical coherence in how that knowledge is represented across subjects, leading to better performance on open-ended questions and cross-subject transfer.

Lu: That emphasis on structural grounding is huge for future AI applications because it means the model isn't just memorizing facts; it's learning the architecture of knowledge itself, which opens up possibilities for much more adaptive learning systems.

Meng: From a practical deployment view, if an AI can understand that algebra requires arithmetic operations as a prerequisite because of the graph structure, that implies it could build tutoring tools that genuinely guide students through the necessary conceptual steps rather than just handing them the final equation.

Lalam: If this K12-KGraph foundation is used to guide training, it means we are moving towards AI systems that possess a deep, structured understanding of how to teach and learn across complex domains.

Tom: So, to wrap up the summary of this paper on "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs," the core thesis is that existing benchmarks fail because they only measure factual recall, and this paper introduces K12-KGraph to provide a curriculum-aligned knowledge graph extracted from official Chinese K–twelve textbooks <ref:2605.09635#pg0,K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational>. It claims this graph closes the gap by providing both a benchmark for probing structural understanding and a training resource for teaching that structural understanding to educational LLMs.

Jane: And the conclusion they draw is that this approach helps us move past simple memorization toward models capable of grasping the organized structure of knowledge, which is essential for effective educational AI.

Lu: The implication I see is that we can start building AI systems not just as answer generators but as genuine knowledge architects within a curriculum context.

Meng: Practically, it means we might see AI tutors that can diagnose conceptual misunderstandings based on where the student's knowledge graph structure breaks down, which is much more useful than a simple right or wrong answer score.

Lalam: For our culture, this work suggests an exciting path toward developing AI that truly understands pedagogy and curriculum design, making educational support systems much more sophisticated.

Conclusion: Tom: So, we've been diving deep into K12-KGraph today, and now we're coming to the close of this segment to talk about what all this means for the field.

Jane: I think it’s really important to focus on that title again—K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs—because it perfectly captures the dual purpose of this work.

Lu: It does, Jane; it shows that we're moving beyond just testing an AI's memory and are building a system to teach it how to structure knowledge itself within a curriculum framework.

Meng: From my side, I see the title highlighting the training aspect as particularly crucial because if you can train an AI on how knowledge is organized rather than just what facts it knows, that opens up much more reliable applications.

Lalam: I feel that the "curriculum-aligned" part is key; it means this isn't just some abstract knowledge graph, but one specifically built from educational materials, which makes its impact very direct on how we design learning tools.

Tom: Exactly, Lalam; and when you consider the authors who put this together, they’ve managed to take a massive challenge—the complexity of structuring K-twelve material across different subjects—and deliver a structured approach that actually works.

Jane: And those authors have done something really neat by creating these benchmarks and training data synthesis methods that are systematically controllable, which is a big step forward for the whole industry.

Lu: That systematic control over difficulty through subgraph analysis is what I find most exciting; it gives us a concrete way to measure exactly where an AI struggles with conceptual relationships rather than just getting the final answer wrong.

Meng: And from an engineering viewpoint, this means we can start building models that aren't just pattern matchers but systems that understand the underlying logic of a subject, which is vital for practical deployment in education.

Lalam: I think the real impact here is on how we develop AI for learning; if the AI understands these explicit relationships, it could build truly personalized tutors that guide students through complex topics by showing them the correct conceptual path.

Tom: That’s a powerful vision, Lalam; and as we wrap up this discussion, it really makes you think about what comes next for educational technology as we start seeing this kind of structured knowledge modeling come to light.

Peking University Institute for Advanced Algorithms Research Zhongguancun Academy OriginHub Technology

cs.CL

Submitted: 2026-05-10

Updated: 2026-10-03

Code: https://github.com/haolpku/K12-Dataset

Project page: https://haolpku.github.io/K12-KGraph-page

Importance score: 90/100

The gist: Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks only measure factual recall, leaving a critical gap in understanding curriculum cognition—the

Key concepts

K12-KGraph
This is a structured knowledge graph built from official Chinese K–12 textbooks, covering subjects like math and science. It connects concepts, skills, figures, and visual elements to map the organized structure of the curriculum.
K12-Bench
A 23,640-question benchmark derived from K12-KGraph. It tests AI models on five types of questions designed to probe structural understanding, such as finding knowledge grounds or tracing prerequisite relationships within the curriculum.
K12-Train Data Synthesis
A training dataset of 7,335 samples created by generating questions from the graph. This data explicitly teaches LLMs how to reason about curriculum structure by asking questions about concepts, skills, and the relationships between them.
Structural Grounding
The advantage gained from using K12-Train data is 'structural grounding.' This means training models to explicitly encode the explicit relationships and organization of knowledge rather than just memorizing isolated facts.

Terminology

Summary

Large language models (LLMs) are increasingly deployed in K–12 education, yet existing benchmarks only measure factual recall, leaving a critical gap in understanding curriculum cognition—the structured knowledge of how concepts are organized and visually presented. This paper introduces K12-KGraph, a curriculum-aligned knowledge graph extracted from official Chinese K–12 textbooks, to close this gap by providing both a benchmark for probing structural understanding and a training resource for teaching it to educational LLMs.

K12-KGraph Construction

The core of the work is K12-KGraph, a heterogeneous property graph covering mathematics, physics, chemistry, and biology across primary through high school. It comprises two components: a textual component capturing curriculum structure (with node types like Concept and Skill) and a multimodal component capturing visual grounding (with Figure and VisualElement nodes). The construction proceeds in five automatic stages:

  1. OCR-based parsing of textbooks into structured Markdown.

  2. A table-of-contents parser to produce a sections index.json manifest, splitting the text into per-section files, and associating images with section text.

  3. LLM-based schema-guided extraction of nodes and edges from each section, including evidence citations or confidence scores for every edge.

  4. Per-section graphs are merged bottom-up: a book-level pass assigns globally unique IDs, deduplicates concepts/skills, and runs depth-first cycle detection on is a and prerequisites for subgraphs to yield valid Directed Acyclic Graphs (DAGs).

  5. Quality control involving manual verification by domain experts to ensure correctness.

K12-Bench Benchmark Construction

From K12-KGraph, the authors derive K12-Bench, a 23,640-question multi-select benchmark designed to probe curriculum cognition through five task families: Ground (Knowledge Grounding), Prereq (Prerequisite Reasoning), Neighbor (Neighbor Recommendation), Evidence (Experiment Evidence Chain), and Locate (Cross-Chapter Indexing). Each task family uses graph-derived templates, with correct answers being true graph neighbors and distractors sampled from structurally proximate but non-answer nodes. This structure ensures that every sample can be traced back to a specific subgraph, making difficulty, coverage, and factual correctness systematically controllable.

K12-Train Data Synthesis

K12-Train is a KG-guided supervised fine-tuning corpus of 7,335 samples designed to explicitly teach curriculum cognition. It is derived through three complementary paths:

  1. Node-grounded QA (LLM-prompted): Generating questions targeting node properties like definitions or formulas for Concepts and Skills.

  2. Edge-grounded QA (LLM-prompted): Creating questions that force the model to articulate the relation itself, such as Why must one learn A before B? for prerequisites for relations.

  3. Exercise-assessment QA (deterministic templates): Using unambiguous edges like tests concept and tests skill to fill templates directly, guaranteeing full factual grounding.

The final dataset is partitioned into K12-Train-Text (2,267 text-only samples) and K12-Train-MM (5,068 multimodal samples), demonstrating the complementarity of textual and visual supervision.

Experimental Results and Findings

Experiments on K12-Bench show that even strong proprietary models like Gemini-3-Flash reach only 57% exact match, indicating a clear gap in curriculum understanding. However, training experiments demonstrate effectiveness: under a strictly matched 2,300-sample SFT budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on both GaokaoBench and EduEval. For VLMs, K12-Train-Full achieves the best overall performance on Gaokao-MM and MDK12-Bench among all compared training configurations. Analysis suggests that the advantage stems from structural grounding (encoding explicit relationships) and pedagogical coherence, leading to gains on open-ended questions and cross-subject transfer across disciplines.

Validation and Quality Assurance

The resources are validated using a tiered strategy centered on the KG foundation. The KG undergoes intensive validation combining automatic structural checks with full human verification by domain-qualified annotators (three per subject). K12-Bench instances are validated via stratified manual sampling, finding 98.4% of sampled items to be fully correct. K12-Train is verified for factual consistency and pedagogical appropriateness; among sampled instances, 96.9% were judged fully correct, confirming that the validated KG provides a reliable basis for both benchmark construction and QA synthesis despite minor surface-level issues in some cases. The work concludes that textual and visual supervision are complementary: K12-Train-Full consistently outperforms both Text Only and MM Only across multimodal benchmarks.

Improvements for AI systems

Here are the specific, actionable improvements an AI system can make by leveraging the K12-KGraph framework:

  1. A capability shift from factual recall to curriculum cognition. The improved system will move beyond simply answering exam questions (like current benchmarks) to understanding the structural logic of knowledge—knowing prerequisites, taxonomies, and conceptual relationships.

  2. Enhanced pedagogical reasoning and tutoring capabilities. The system can now act as a teacher that explains the why behind a concept (e.g., You must learn arithmetic before linear equations because of the prerequisite relationship). This allows for more effective student support in K-12 settings than current LLMs allow.

  3. Improved performance on complex, structural reasoning tasks. The system will show superior ability in tasks requiring graph traversal and structural inference, such as:

  4. Advanced Prerequisite Reasoning: Accurately determining the necessary sequence of learning (e.g., To solve this problem, you first need to master X and then Y).

  5. Neighbor Recommendation/Concept Discovery: Suggesting related concepts or skills based on direct curriculum connections (e.g., Since you are studying 'linear equations,' you should also review 'combining like terms').

  6. Robust Multimodal Reasoning for STEM subjects. The system can analyze diagrams, experimental setups, and figures in textbooks to extract visual evidence and link it directly to textual concepts (e.g., identifying which specific diagram explains a chemical reaction or what an apparatus is used for in a physics lab).

  7. Superior Domain-Specific Fine-Tuning Efficiency (Sample Efficiency). The system can achieve state-of-the-art performance on Chinese K–12 educational tasks using significantly fewer, highly curated samples (e.g., 2,300 samples) compared to training on massive general instruction datasets, making domain adaptation faster and more cost-effective.

  8. Complementary Textual and Visual Grounding for Vision-Language Models (VLMs). When used with vision models, the system can leverage the combined supervision from textual structure and visual evidence to achieve better performance across multimodal benchmarks like Gaokao-MM or K12Vista, as it learns both what concepts are and how they are visually represented.

  9. Increased Reliability through Structural Grounding. Because benchmark questions (K12-Bench) and training data (K12-Train) are derived deterministically from a verified Knowledge Graph, the system's performance is directly traceable to the underlying textbook structure, leading to higher factual correctness and reduced hallucination in educational contexts.

Sources

Related papers