Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings".
Jane: The paper was written by Yahya Badran and Christine Preisach from University of Applied Sciences Karlsruhe and Pädagogische Hochschule Karlsruhe.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the channel, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings." Jane, I've got to say, the title alone had me intrigued.
Jane: Oh, absolutely, Tom. And I think for our listeners who might not be deep in the education tech world, we should break that down. Knowledge tracing is basically the science of predicting whether a student will get the next question right, based on everything they've done before. It's how intelligent tutoring systems decide what to show you next.
Tom: Right, and the authors here are Yahya Badran and Christine Preisach, from the University of Applied Sciences Karlsruhe and the Pedagogical University in Karlsruhe. They're tackling a really sneaky problem. You know how sometimes a math problem might test two skills at once, like addition and subtraction?
Jane: Exactly. So the system might log that as two separate entries in the student's history, both with the same correct or incorrect label. And that's where the trouble starts. The model can cheat. When it's trying to predict the outcome for the second skill, it can just peek at the label from the first skill, because they came from the same question.
Tom: And that's the "leakage" in the title. The model isn't actually learning the student's knowledge state; it's just learning to copy the answer from a neighboring entry. That's a huge problem because in the real world, when a student is facing a new question, there's no label to peek at. So the model's performance in the lab is artificially inflated.
Jane: It's like studying for a test with the answer key accidentally printed on the back of the page. You'd score great on the practice test, but you'd fail the real one. The authors are saying, look, our models are doing exactly that, and we need to stop it.
Tom: And they don't just point out the problem. They propose a fix. They introduce a special "MASK" label that replaces the ground-truth answer in the input, so the model literally cannot see the leaked information. It's a clever, simple idea. I'm really looking forward to seeing how they implemented it in the actual models.
Jane: Me too. And there's a second part to the title, "Recency-Aware Embeddings." That's about telling the model how long it's been since a student last practiced a specific skill. That's crucial for modeling forgetting, which is a huge part of learning. I think this paper has the potential to make a lot of existing models much more honest and much more accurate.
Tom: And that's what we're here to dig into. So stick around, because next we're going to look at the core problem they're solving and why it's so widespread in the field.
Summary and Core Problem: Tom: So, Jane, we're back with "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings," and I want to get into the nitty-gritty of this leakage problem. It's not just a minor bug; it's baked into how a lot of these models are trained.
Jane: Right. The paper explains that many popular models, like DKT and AKT, rely on expanding each question into its individual knowledge concepts, or KCs. So a question about "four minus two plus three" becomes two separate interactions: one for "subtraction" and one for "addition."
Tom: And because both of those interactions come from the same student answer, they share the same label. The model sees that shared label and learns a shortcut. The authors actually did a really neat experiment to show this. They created a synthetic dataset called CorrAS09, where they artificially duplicated every knowledge concept.
Jane: So every question suddenly had two highly correlated KCs. And what happened? The performance of the standard DKT model on that corrupted dataset was significantly worse than on the original, clean dataset. That's the leakage hurting the model's ability to generalize. It was learning to rely on the duplicate label rather than the student's actual history.
Tom: And they also showed that during training, the validation scores on these leaky datasets were unrealistically high. It looked like the model was doing great, but it was actually just memorizing the correlation. The authors call this out as a serious problem for model selection, because you might pick a model based on inflated scores that will never hold up in the real world.
Jane: Exactly. And that's why their solution is so important. They propose the "Mask Label" method. In the expanded sequence, for any KC that isn't the last one in its question group, they replace the ground-truth label with a special MASK token. So the model only sees the actual response for the final KC of that question.
Tom: That way, when the model is processing the first KC, it has no idea what the correct answer was. It has to rely on the student's past performance. And this isn't just a training trick. They apply it during inference too, so the model is always working with the same kind of information it would have in production.
Jane: It's a beautifully simple fix. You're not changing the model architecture at all. You're just changing how you feed the data in. That means it can be applied to DKT, AKT, SAKT, basically any model that uses this KC expansion approach.
Tom: And they didn't stop there. They also introduced the recency encoding, which we mentioned earlier. But I think we should save that for the next segment, because that's where they talk about the actual improvements and results.
Jane: Good call. Let's dig into the results next.
Improvements and Results: Tom: Welcome back. We're still on "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings." So Jane, we've established the leakage problem and the MASK label fix. Now, what about this recency encoding part? What did they actually do?
Jane: So, the idea is to explicitly tell the model how many steps it's been since the student last saw a particular knowledge concept. Think about learning a new language. If you practiced a word yesterday, you'll probably remember it. If you practiced it three weeks ago, you might have forgotten it. That distance matters.
Tom: And they encode this distance using something called learnable Fourier features. It's a fancy way of saying they create a numerical representation of that time gap that the model can learn to interpret. They add this to the input embeddings of models like AKT and SAKT.
Jane: And the results are pretty impressive. They tested this on several benchmarks, including ASSISTments2009 and Duolingo2018. On the Duolingo dataset, which is all about language learning and has a high number of KCs per question, their DKT model with both the MASK label and recency encoding jumped from an AUC of about zero point six five to a whopping zero point eight nine.
Tom: That's a massive jump. And it's not just DKT. Their best model overall was AKT with both improvements, which they call AKT-MLd. It consistently beat the original AKT on every single benchmark they tested.
Jane: Right. And they also compared their recency encoding to the standard positional encoding used in transformer models. Their recency encoding performed better. That makes sense, because positional encoding just tells you where you are in the sequence, but recency encoding tells you something meaningful about the learning process itself.
Tom: And I love that they tested all this against models that don't have the leakage problem, like QIKT and DKVMN. Their modified models, especially AKT-MLd, were competitive with or even better than those leakage-resistant baselines. So they're not just fixing a bug; they're actually pushing the state of the art forward.
Jane: Exactly. And the beauty of it is that these are purely embedding-level changes. They're not inventing a new neural network architecture. They're just making the inputs smarter. That means the computational cost is minimal, and it's easy for other researchers to adopt these techniques in their own models.
Tom: So we have a fix for a widespread problem and a new way to encode temporal information, both of which are simple to implement. I think we need to bring in our senior researcher, Lu, to talk about the bigger picture here.
Conclusion: Tom: So, Lu, we've been talking about the practical results of "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings." What do you think the broader impact of this is going to be?
Lu: I think the most important thing is that it forces the field to be more honest. For years, we've been comparing models on benchmarks that had this hidden leakage. Some models might have looked better simply because they were better at exploiting the leak, not because they actually understood the student. This paper gives us a clean way to level the playing field.
Jane: That's a really good point, Lu. It's not just about making one model better; it's about making the entire evaluation process more trustworthy. If we're going to deploy these systems in real classrooms, we need to know that their performance is real.
Meng: And from an engineering standpoint, I appreciate that they didn't ask us to build a whole new infrastructure. The MASK label is just a new token in the embedding table. The recency encoding is a small feed-forward network. You can drop these into an existing pipeline without rewriting everything.
Tom: And that's huge for adoption. If a school district is already using a system based on DKT, they don't need to buy a new system. They can just update the software.
Lalam: If I may add, the cultural impact is significant. Intelligent tutoring systems are becoming more common, especially in regions with teacher shortages. If we can make these systems more accurate and more honest, we can provide better personalized education to more students. The recency encoding, in particular, helps the system adapt to the natural human process of forgetting, which makes the interaction feel more like a real tutor who remembers what you've struggled with.
Jane: That's a beautiful way to put it, Lalam. The system isn't just a static database of questions; it's actively modeling the student's memory.
Tom: Alright, so let's wrap this up. We've got a paper that identifies a critical flaw in knowledge tracing, provides a simple fix, and adds a new way to model learning dynamics. The results are clear, and the methods are easy to adopt.
Lu: And it opens up a lot of future work. Now that we have leakage-free benchmarks, we can start comparing models on a level playing field. And the recency encoding could be extended to include actual time, like minutes or hours, rather than just the number of steps.
Meng: I'd also like to see how these embeddings interact with other features, like response time or hint usage. There's a lot of room to build on this.
Tom: Well said. We've been discussing "Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings" by Badran and Preisach. A fantastic piece of work that makes AI in education a little more reliable and a lot more human-aware. Thanks for joining us, everyone. We'll see you on the next one.
Yahya Badran, Christine Preisach
University of Applied Sciences Karlsruhe · Pädagogische Hochschule Karlsruhe
cs.CY, cs.AI, cs.LG
Submitted: 2026-08-09
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 73/100
Terminology
Summary
Summary
This paper addresses two key limitations in deep learning-based Knowledge Tracing (KT) models: label leakage and the lack of explicit recency information. KT models predict a student’s future performance based on their sequence of interactions with learning content, often leveraging Knowledge Concepts (KCs) to mitigate data sparsity. However, when a question involves multiple KCs, expanding the question-level interaction into KC-level interactions introduces label leakage, where the model can exploit the shared ground-truth label of KCs from the same question to inflate its predictions artificially. The paper states: This occurs when models inadvertently exploit shared labels among KCs originating from the same question to infer the ground-truth label.
This leakage is problematic because ground-truth labels are unavailable during deployment, label leakage can artificially inflate evaluation metrics.
To prevent label leakage, the authors propose a Mask Label
method. They introduce a special MASK label that replaces ground-truth labels in cases where leakage is likely, particularly when multiple KCs belong to the same question. The paper explains: Any ground-truth response preceding another KC within the same question is replaced with MASK, explicitly preventing intra-question label leakage.
For a question q associated with KCs (c1, c2, c3) and response r, the expanded sequence becomes: ..., (q, c1, MASK), (q, c2, MASK), (q, c3, r),...
Only the final KC retains its actual ground-truth label. The MASK token is applied exclusively to input embeddings, while output predictions remain constrained to labels 0 and 1. This procedure is applied during both training and inference, eliminating the need for special-case handling at inference time. Models adopting this method are denoted with the suffix -ML
(e.g., DKT-ML, DKT-ML+, AKT-ML, SAKT-ML).
In addition, the paper introduces Recency Encoding,
which explicitly encodes the step-wise distance between the current item and its most recent previous occurrence. The authors argue: This distance is important for modeling learning dynamics such as forgetting, which is a fundamental aspect of human learning, yet it is often overlooked in existing models.
The encoding is based on learnable Fourier features, defined as: γ(d) = [cos(dwf + bf), sin(dwf + bf)], where wf and bf are trainable parameters. The resulting features are passed through a feedforward transformation with a GeLU activation. Models incorporating this encoding are denoted with a superscript d (e.g., AKT-MLd, SAKT-MLd, DKT-MLd). The authors note that Recency Encoding demonstrates improved performance over traditional positional encodings on multiple KT benchmarks.
The paper evaluates several alternative mitigation strategies, including Averaged Embeddings (DKT-Fuse, AKT-Fuse), Autoregressive Decoding (DKT-AD), and Special Attention Masking (AKT-QM). However, the Mask Label method is shown to be superior: MASK-based models (e.g., DKT-ML and AKT-ML) outperform other mitigation strategies across most datasets.
The authors also highlight that the Mask Label method is simpler to implement, computationally cheaper, and broadly applicable across various architectures.
Experiments are conducted on four benchmark datasets: ASSISTments2009, Algebra2005, Riiid2020, and Duolingo2018, plus a synthetic variant CorrAS09 designed to amplify label leakage. The results consistently show that the proposed methods improve prediction accuracy. For instance, on ASSISTments2009, AKT-ML achieves an AUC of 0.7543 compared to 0.7334 for the original AKT. On Riiid2020, AKT-ML achieves 0.7411 versus 0.6136 for AKT. The best overall performance is achieved by AKT-MLd, which incorporates both mask labels and recency encoding, achieving AUC scores of 0.7566, 0.8325, 0.7441, and 0.8947 on ASSISTments2009, Algebra2005, Riiid2020, and Duolingo2018, respectively.
The paper also demonstrates that label leakage inflates validation scores during training, particularly on datasets with higher numbers of KCs per question. The authors show that DKT achieves unrealistically high validation scores on CorrAS09 and Algebra2005,
whereas DKT-ML yields consistent validation performance across datasets. This underscores the importance of using leakage-aware strategies throughout the training pipeline.
The authors conclude: "Our findings show that addressing label leakage and incorporating recency information can substantially improve KT performance with only minimal architectural changes. This makes the proposed methods practical and broadly applicable to a wide range of KT models." The contributions are summarized as: (1) empirically demonstrating and quantifying label leakage in widely used KT models, (2) proposing a simple and computationally efficient method to prevent label leakage using masked ground-truth labels, (3) introducing recency encoding that explicitly encodes the number of steps since each item’s last occurrence, and (4) performing extensive experiments showing improvements across multiple KT models and benchmark datasets.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Implementation: Add a masking step before embedding construction that replaces ground-truth labels with a special MASK token for all KCs belonging to the same question except the last one.
What the improved system can do:
-
Prevent the model from exploiting intra-question label correlations during training
-
Eliminate artificially inflated validation scores (e.g., DKT on CorrAS09 improves from AUC 0.6312 to 0.7163)
-
Maintain consistent performance between training and deployment conditions
Sources
Related papers
- Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus
- Generative AI Purpose-built for Social and Mental Health: A Real-World Pilot
- PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
- What is an intelligent system?
- AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study
- Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework