T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation

arXiv:2506.12075 · cs.IR, cs.AI · Submitted 2025-06-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation".

Jane: The paper was written by Nirmal Gelal, Chloe Snow, Ambyr Rios, Kathleen M. Jagodnik and Hande Küçüük McGinty from Department of Computer Science and Department of Curriculum and Instruction and Kansas State University, Manhattan, Kansas, United States.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Jane: Okay, so we’ve established that the goal is making high school literature selection easier, using "T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation." Now, the paper summary dives into *how* they achieve this.

Tom: What I took away was that they aren't just relying on simple keyword matching; they are leveraging a knowledge graph structure to understand relationships between texts, concepts, and even characters.

Jane: That’s the core mechanism: instead of just saying "read Book A because it mentions Topic X," the system understands *why* Book A is relevant to Topic X within the context of a specific curriculum goal.

Lu: The knowledge graph acts as a sophisticated semantic layer, defining not just that two things are related, but *how* they are related—is it cause-and-effect? Is it thematic influence?

Meng: And this makes the recommendations far more robust than what standard search engines offer because they're factoring in multiple layers of context simultaneously.

Lalam: It sounds like the AI is moving beyond simple information retrieval and into deep contextual understanding, which is a massive leap for educational technology.

Tom: So, if I understand correctly, the system builds a map where every node is an entity—a character, a theme—and every edge shows the relationship between them.

Jane: Exactly. And when a teacher inputs a learning objective, the knowledge graph helps pinpoint specific textual segments that best fulfill that objective across multiple sources.

Lu: I'm thinking about the scalability here; if you feed it enough diverse data, it could map out entire cultural epochs, showing the intellectual connections between seemingly unrelated authors.

Meng: The practical challenge would be populating that graph initially—you need expert human input to define those relationships accurately before the AI can even start recommending.

Lalam: But once that foundational knowledge is established, the system has the potential to democratize access to deep academic resources by guiding users through complexity.

Tom: It’s really about making complex information digestible without losing its depth, which is a delicate balance. We should talk more about what improvements they suggest next.

Improvements: Jane: Last time we were talking about the summary section of "T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation." If the summary showed us *what* it does, this segment talks about how they improve it.

Tom: It seems like the authors are addressing limitations by suggesting ways to make the recommendation process even more fine-grained and adaptable for different teaching styles.

Jane: They're moving past just "this text is good" to suggesting *why* it's good and *how* a teacher should use it in a lesson plan.

Lu: The improvements they suggest point toward making the scaffolding dynamic—meaning the system adapts its recommendations as the student actually interacts with the material, rather than just giving one static list.

Meng: That interactivity is key; an ideal engineering solution wouldn't just spit out links, it would suggest specific passages and even prompts for discussion based on those passages.

Lalam: This moves the technology closer to being a co-pilot for learning, where the AI is actively participating in the pedagogical process alongside the human teacher.

Tom: So, they’re not just recommending; they’re designing micro-learning experiences right inside the system. It's curriculum design powered by AI.

Jane: That would save teachers enormous amounts of time because they wouldn't have to manually sift through hundreds of pages to find the perfect example for a specific discussion point.

Lu: I can envision this expanding into personalized learning paths, where if a student struggles with character motivation, the system automatically recommends texts that specifically focus on unreliable narration or internal conflict.

Meng: Implementing those adaptive prompts would require integrating natural language generation capabilities right into the recommendation engine itself—that's a substantial layer of complexity.

Lalam: The impact here is creating educational equity; students in under-resourced schools could gain access to sophisticated, personalized learning tools that were previously only available in top academic institutions.

Tom: It feels like they are solving the problem of information overload for both the student and the teacher at the same time. Let's wrap up our discussion by looking at the big picture impact.

Conclusion: Tom: So, we've covered how "T-TExTS (Teaching Text Expansion for

Conclusion: Tom: So we’ve spent a lot of time breaking down how T-TExTS uses knowledge graphs to help teachers select literature, but we need to bring it all together now.

Jane: It’s clear that the system provides a way for educators to create diverse, thematically aligned text sets without having to manually search through massive databases.

Lu: And I think that's where the real potential lies; we are building a tool that understands deep semantic connections rather than just finding surface-level keywords.

Meng: From an engineering viewpoint, it’ is a highly scalable architecture because the knowledge graph structure is inherently robust even when you have hundreds of thousands of possible texts.

Lalam: The biggest cultural impact I see is that this democratizes access to complex literature, making sophisticated teaching resources available to anyone looking for them.

Tom: That's exactly right, Jane; it’ gives teachers a powerful scaffolding tool for informed curricular decisions.

Jane: It seems like we found a great balance between algorithmic tuning and expert-guided input in the T-TExTS system.

Lu: The consistency of the results across different dataset sizes is incredibly encouraging because it suggests this approach generalizes well as more content becomes available.

Meng: I agree with Lu; it’ proves that this methodology is practical, not just for a small set of books, but for a whole curriculum.

Lalam: It shows that our AI can actually grasp pedagogical intent rather than just mimicking user behavior, which is a huge step forward in human-centered design.

Tom: It’s amazing to see the full potential of T-TExTS, as it solves the real-world problem of finding diverse texts for the English Literature curriculum.

Jane: We hope that this system provides significant relief and support to teachers across all grade levels.

Lu: I'm looking forward to seeing how these kinds of knowledge-driven systems interact with more dynamic learning environments.

Meng: It’s a practical solution, and it gives us a very solid foundation for future work in curriculum design tools.

Lalam: This is the kind of innovation that allows cultural understanding to grow alongside the education of its students.

cs.IR, cs.AI

Submitted: 2025-06-06

Updated: 2026-05-13

Code: https://github.com/koncordantlab/TTExTS

Importance score: 81/100

The gist: This paper presents T-TExTS (Teaching Text Expansion for Teacher Scaffolding), a knowledge graph (KG)-based recommendation system designed to assist high school English Literature teachers in

Key concepts

Knowledge Graph
A structured map where every entity (like a character or theme) is a node, and every relationship between them is an edge. This allows the system to understand complex connections, such as cause-and-effect or thematic influence, rather than just surface-level keywords.
T-TExTS
An acronym for Teaching Text Expansion for Teacher Scaffolding. It is a system designed to improve literature selection by using a knowledge graph structure. The goal is to help teachers find specific textual segments that best fulfill a defined learning objective.
Semantic Understanding
The AI's ability to understand the deeper meaning and context of information, not just retrieve keywords. This allows the systems to identify how texts relate to curriculum goals, making recommendations far more robust than standard search engines.

Terminology

Summary

This paper presents T-TExTS (Teaching Text Expansion for Teacher Scaffolding), a knowledge graph (KG)-based recommendation system designed to assist high school English Literature teachers in assembling diverse, thematically aligned text sets. By utilizing a pedagogy-first ontology, the system aims to ease the burden of text selection by suggesting works based on pedagogical merit rather than surface-level metadata, thereby supporting more informed and inclusive curricular decisions.

Ontology and Knowledge Graph Construction

The researchers constructed a domain-specific ontology using the Knowledge Acquisition and Representation Methodology (KNARM) to transform qualitative domain knowledge into a machine-interpretable representation. This process involved several stages, including sub-language analysis, unstructured interviews with literacy teachers, and the creation of metadata frameworks. The resulting knowledge graph was instantiated with two distinct components:

  1. A Terminological Box (TBox) that defines class hierarchies and properties such as Text, Author, Genre, Theme, and TextComplexity.

  2. An Assertional Box (ABox) that records specific facts about individuals through relations like has author, has genre, and has theme.

To ensure accuracy, texts were evaluated using both quantitative tools—such as the Lexile Analyzer and Flesch-Kincaid Grade Level Calculator—and qualitative rubrics to capture instructional merit descriptors. This allowed the system to represent the interconnected nature of literary concepts rather than relying on simple keyword overlap.

Graph Embedding Strategies

To project symbolic relationships into a continuous vector space, T-TExTS employs various random-walk-based embedding methods. The study evaluates four specific strategies to determine how they quantify the pedagogical proximity between texts:

  • DeepWalk, which utilizes uniform random walks to explore neighborhoods.

  • A biased random walk that uses expert-guided transition probabilities to emphasize pedagogically salient relations.

  • Node2Vec, a parameterized random walk that uses return (p) and in-out (q) parameters to balance Breadth-First Search (BFS) and Depth-First Search (DFS).

  • A hybrid embedding that combines structural and pedagogical signals through embedding concatenation.

These methods allow the system to identify recommendations that are thematically and contextually aligned with a teacher’s anchor text based on graph structure.

Experimental Performance and Scaling

The system was tested across three dataset configurations of increasing scale (98, 196, and 351 texts). The results indicate that while ranking metrics like Hits@K, MRR, and nDCG naturally decline as the candidate pool expands, the AUC remains remarkably stable across all scales. A key finding is that traversal-level expert weighting alone does not outperform algorithmic structural tuning. Specifically, Node2Vec achieves the highest Area Under the Curve (AUC) at every dataset size (0.9642–0.9750) and the strongest ranking metrics at larger scales. However, the hybrid model remains a highly competitive practical deployment choice, as it maintains a high AUC while remaining within a few percentage points of Node2Vec on every ranking metric.

Qualitative Case Analysis

A qualitative evaluation using the anchor text 1984 illustrated how different strategies surface different but pedagogically defensible associations. The analysis revealed several important patterns:

  1. All configurations successfully ranked Fahrenheit 451 first, demonstrating a strong structural alignment in the knowledge graph regarding shared genre and themes.

  2. Node2Vec was the only configuration to recover all five expert-curated ground-truth texts, successfully surfacing structurally distant but thematically aligned works like Marrow Thieves.

  3. The system consistently produced strong near-miss recommendations, such as The Giver and Scythe, which represent legitimate alternatives that a teacher might consider when designing a curriculum.

Improvements for AI systems

1. Hybrid Embedding Concatenation (Biased + Uniform Random Walks)

  • What the improved AI can do: The system can provide high-precision, expert-guided recommendations (e.g., prioritizing specific genres or themes) without sacrificing global topological accuracy. By concatenating biased embeddings with uniform ones, the AI maintains high AUC for link prediction while simultaneously providing the interpretability and precision required for specialized domain constraints.

2. High- q Parameterized Node2Vec Traversal

  • What the improved AI can do: The system can perform deep-first search (DFS)-biased exploration within knowledge graphs to discover thematically distant but contextually aligned entities. This allows the AI to break out of local neighborhood clusters (which often lead to redundant, surface-level recommendations) and identify non-linear semantic connections across a graph.

3. Multi-Layered Neuro-Symbolic Ontology Integration

  • What the improved AI can do: The system can move beyond surface-level metadata (keywords/popularity) to recommend items based on functional utility or instructional merit. By modeling qualitative, non-linear attributes (e.g., complexity levels, levels of meaning, and pedagogical scaffolding) within a formal TBox/ABox structure, the AI can suggest content that fulfills specific structural or developmental requirements rather than just content similarity.

4. Ontological Near-Miss Evaluation Metrics

  • What the improved AI can do: The system can be trained and validated using metrics that reward pedagogical equivalence classes rather than binary ground-truth matching. This enables the AI to recognize and prioritize alternative valid candidates—items that are structurally distant in the graph but ontologically similar—preventing the model from being penalized during training for suggesting legitimate, high-quality alternatives.

5. Decoupled TBox/ABox Scaling Architecture

  • What the improved AI can do: The system can scale its entity and relationship density (ABox) indefinitely while maintaining a stable, high-fidelity schema (TBox). This ensures that embedding quality and AUC remain stable as the knowledge graph grows, preventing the degradation of recommendation accuracy typically seen when expanding specialized, expert-curated datasets.

Sources

Related papers