Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

arXiv:2509.21514 · cs.LG, cs.CL · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing".

Jane: The paper was written by Joshua Mitton, Prarthana Bhattacharyya, Ralph Abboud and Simon Woodhead from Eedi and Renaissance Philanthropy.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, we’ve got a paper today that I think is going to get people talking — it’s called "Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing." And honestly, the title alone tells you exactly what it’s about. These knowledge tracing models, the ones that predict how students will do on the next math question, they’re being used in real classrooms now. But they just spit out a prediction with no sense of when they might be wrong.

Jane: Right, Tom, and that’s the gap this paper goes after. The authors — Mitton, Bhattacharyya, Abboud, and Woodhead from Eedi — they’re saying, look, we don’t just want a model that’s accurate on average. We want a model that knows when it’s out of its depth. So they built a layer on top of existing models that lets the model say, "I’m not confident about this one, let a human teacher handle it."

Tom: And that’s the "defer" part of the title. Instead of forcing the model to answer every single question, you let it abstain on the ones it’s most unsure about. The paper shows that if you defer just twenty percent of the most uncertain predictions, accuracy jumps by two to three percentage points across three different model architectures — DKT, SAKT, and AKT. No retraining, no new architecture. Just a smarter way to use what’s already there.

Jane: What I love is how they measure that uncertainty. They use something called Monte Carlo Dropout. During training, neural networks randomly turn off some neurons to avoid overfitting. Normally at test time you turn that off. But here, they keep it on and run the model a hundred times, and if the predictions bounce around a lot, that’s the model telling you it’s not sure.

Tom: It’s like asking a student the same question a hundred times and seeing if they give the same answer. If they’re all over the place, you know they don’t really know it. And the paper shows this signal is genuinely useful — the predictions they defer have an error rate one and a half to one point six times higher than the ones they keep. So they’re really finding the mistakes.

Jane: And that’s the responsible part. It’s not just about accuracy for accuracy’s sake. It’s about building systems that know their own limits. I’m really curious to see how they checked whether this deferral is fair across different students — that’s where we’re headed next.

Tom: Stay with us, because that fairness analysis is genuinely surprising. The weakest students actually get deferred less, not more. We’ll dig into that right after the break.

Summary: Tom: Welcome back. We’re talking about "Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing," and Jane, you teased the fairness angle. Let’s get into it.

Jane: So the paper checks whether the model’s abstention is fair across student ability levels. They split students into four quartiles by ability, and they found that the deferral rate stays between sixteen and twenty-three percent across all four groups. That’s close to the twenty percent uniform baseline. But here’s the kicker — the weakest students, the bottom quartile, actually get deferred less, around sixteen to seventeen percent.

Tom: That’s the opposite of what you’d fear. You’d worry that a model might just punt on struggling students and hand them off to teachers all the time. But this model doesn’t do that. It’s actually more confident about those students, which makes sense because their performance is more predictable — they tend to get things wrong consistently.

Jane: And it’s not just about student ability. They also check question difficulty. They split questions into difficulty quartiles, and the deferral is targeted within every single quartile. The deferred set always has a higher error rate than the kept set. The effect is actually strongest on the hardest questions — there, the error rate in the deferred set is two point three to two point five times higher than in the kept set.

Tom: So the model is really good at saying, "This hard question, I have no idea what this particular student will do." And that’s exactly when you want a human in the loop. But here’s the thing that really got me — they compare their uncertainty signal against a much simpler baseline. A classic psychometric model called Item Response Theory, or IRT.

Jane: And that comparison is brutal. The IRT baseline, which uses question difficulty and student ability to estimate confidence, gives an AUC lift of about zero point four percentage points at eighty percent coverage. The Monte Carlo Dropout signal gives two to two point four percentage points. That’s roughly five times better.

Tom: Which raises the obvious question — why? Why does a simple, well-calibrated statistical model do so much worse than running a neural network a hundred times with dropout?

Jane: That’s exactly what the paper investigates next. They do this variance decomposition thing to figure out what the uncertainty signal is actually tracking. And the answer is surprising — it’s not tracking question difficulty, it’s not tracking student ability, it’s not even tracking whether the student has seen similar content before. All of those classical factors together explain less than four percent of the variance in the model’s uncertainty.

Tom: Four percent. That’s almost nothing. So the model is uncertain about something that the classical features just can’t see. That’s a really deep finding, and I want to know what that something is.

Jane: That’s the next segment — they try to dig into what’s left, and it’s a wild ride. Stick around.

Improvements: Tom: We’re back with "Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing," and Jane, you left us hanging on that four percent. What’s the model actually uncertain about?

Jane: So the paper does this clever thing. They take the model’s epistemic uncertainty — that’s the part that comes from not knowing the right answer, not from random noise — and they try to predict it using all the classical features. Question difficulty, student ability, IRT-style outcome ambiguity, and even curriculum coverage at four different levels of granularity. Under linear modeling, all of that together explains less than four percent.

Tom: But then they get more aggressive. They throw a random forest at it, a non-linear model, and even then the explained variance only goes up to about ten percent for SAKT and AKT, and twenty-three percent for DKT. So even with a much more flexible model, seventy-seven to ninety percent of the signal remains unexplained.

Jane: And that’s the core argument of the paper. The uncertainty the model feels is architecture-specific. It’s about the model’s own internal representations of a student’s unique sequence of answers. Two students with the same overall ability and the same question difficulty can produce wildly different uncertainty levels because the model has learned something specific about their trajectory.

Tom: That’s the improvement over the simpler baselines. The paper is essentially saying, you cannot approximate this with a spreadsheet. You cannot approximate it with IRT. You have to actually run the model and see where it’s uncertain. That’s the only way to get this signal.

Jane: And that has real implications for deployment. If you’re building a tutoring system, you don’t want to just use a heuristic like "if the student got the last three questions wrong, flag them for review." That misses most of the signal. The paper shows that a model-native uncertainty probe is necessary.

Tom: I want to bring in Lu and Meng here, because I think they’ll have strong opinions on this. Lu, you work with these models every day — does this match your experience?

Lu: It does, Tom. And I think the most striking part is the BALD decomposition they use. They separate aleatoric uncertainty — the irreducible randomness — from epistemic uncertainty — the reducible uncertainty about the model’s own parameters. And the fact that classical features explain so little of the epistemic part tells me that the model is encoding something about student learning dynamics that we don’t have a name for yet.

Meng: From an engineering standpoint, though, I want to know about cost. Running a model a hundred times at inference is expensive. For a real-time tutoring system, that could be a problem.

Jane: That’s a fair concern, Meng. The paper doesn’t address latency directly, but a hundred forward passes on a small model like DKT, which is only one point two million parameters, is not that bad. You could probably get away with fewer samples — the paper uses a hundred, but the variance signal might stabilize well before that.

Tom: And the payoff is real. You’re getting a two to three point accuracy boost just by knowing when to abstain. That’s the kind of improvement that matters in a classroom where a wrong prediction can send a student down the wrong learning path.

Lu: I’d push further. The paper says this is complementary to fairness audits and classroom evaluation, not a replacement. But I think the deferral signal itself could be used to build better adaptive learning systems — ones that know when to challenge a student and when to back off.

Meng: And I’d want to see this tested in a real classroom before trusting it. The paper is honest about that limitation — they haven’t run an A/B test with actual students and teachers.

Jane: That’s the honest caveat. The metrics improve, but we don’t yet know if that translates to better learning outcomes. That’s the next big question, and it’s where we’re heading in the conclusion.

Conclusion: Tom: We’re wrapping up our discussion of "Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing," and Jane, I think we’ve landed on something important.

Jane: We have. The paper’s central message is that responsible AI deployment isn’t just about making models more accurate — it’s about making them honest about their own limits. By using Monte Carlo Dropout to quantify uncertainty, they’ve shown that existing knowledge tracing models can abstain on their most uncertain predictions and get a meaningful accuracy boost without any retraining.

Tom: And the fairness piece is what really sold me. The deferral isn’t dumping struggling students on teachers — it’s actually more conservative with them. And the targeting holds across every question difficulty level. That’s a system you can trust.

Jane: The deeper finding, though, is that the uncertainty signal is mostly architecture-specific. The classical psychometric features — difficulty, ability, coverage — explain almost nothing. That means you can’t shortcut this with a simpler model. You need the model’s own internal signal.

Lu: And that’s the part I’ll remember. It suggests that these models are tracking something about student learning that we don’t yet have a language for. That’s a research opportunity.

Meng: From my side, the practical takeaway is clear. If you’re deploying a knowledge tracing model, you should build in this deferral mechanism. The cost is low, the benefit is real, and it’s the responsible thing to do.

Tom: The paper is honest about its limits — single dataset, no classroom trial yet, and the IRT comparison is just one baseline. But as a proof of concept, it’s compelling. It shows that selective prediction with model-native uncertainty is a necessary part of responsible deployment.

Jane: And with that, we’ll say goodbye to "Knowing When to Defer." It’s a paper that gives us a concrete tool for making AI systems safer in education, and it opens up a whole line of research into what models are actually uncertain about. Thanks for listening, and we’ll see you next time.

Tom: Take care, everyone.

Joshua Mitton, Prarthana Bhattacharyya, Ralph Abboud, Simon Woodhead

Eedi · Renaissance Philanthropy

cs.LG, cs.CL

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 10 pages, 7 figures. Joshua Mitton and Prarthana Bhattacharyya contributed equally to this paper

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

Key concepts

Knowledge Tracing Models
These models predict how students will perform on future questions, such as in math. They are currently used in real classrooms to assess student progress and guide learning paths.
Monte Carlo Dropout
A technique used to measure a neural network's uncertainty. By running the model many times with random neuron deactivation (dropout), researchers see if the predictions vary widely, indicating low confidence.
Deferral
The process of having an AI model intentionally abstain from answering a question when it is most unsure about the prediction. This selective approach improves overall accuracy and safety.
Epistemic Uncertainty
A type of uncertainty that comes from the model's lack of knowledge or parameters, as opposed to random noise. The paper suggests this signal is unique to the model's internal representations.

Terminology

Summary

Summary

This paper addresses the problem of responsible real-world deployment of Knowledge Tracing (KT) models, which predict how a student will respond to future questions based on past interactions. The authors note that while KT models have evolved from Bayesian approaches to LSTMs, self-attention, and graph-based methods, actual performance gains in accuracy are often marginal and existing KT models typically struggle with low specificity, which allows incorrect student answers to slip by undetected during real-world deployment.

The core contribution is treating responsible KT deployment as a selective prediction problem. The authors introduce an intrinsic selective prediction layer designed to integrate directly with existing KT models, which calculates its own uncertainty, abstains from making the most unreliable predictions, and gracefully defers those specific cases to a human teacher. They use Monte Carlo Dropout (MC-Dropout) to quantify uncertainty, applied across three KT architectures: DKT (Deep Knowledge Tracing, an LSTM-based model), SAKT (Self-Attentive Knowledge Tracing), and AKT (Attentive Knowledge Tracing). The analysis is based on the Eedi mathematics dataset, with a training split of 11,994 students and a validation split of 1,493 students, each with 100 question-answer pairs.

The paper offers four primary contributions:

1. Selective prediction improves metrics without retraining. At an operating point of 80% coverage (deferring the 20% most uncertain predictions), MC-Dropout variance increases accuracy by 2.3 to 3.0 percentage points, AUC by 1.9 to 2.4 percentage points, and F1 by 1.4 to 4.3 percentage points across all three architectures, with all gains statistically significant (95% bootstrap confidence intervals never cross zero). The abstention is highly targeted: the deferred set exhibits an error rate 1.45 to 1.60 times higher than the kept set. This targeting holds true within every quartile of question difficulty (with the strongest targeting on the hardest questions, showing 2.3 to 2.5 times higher error rates) and remains equitable across different student ability levels (abstention rates range between 16% and 23% across all four student-ability quartiles, with the weakest students receiving the lowest abstention rate of 16–17%).

2. MC-Dropout outperforms a calibrated IRT baseline. The authors compare MC-Dropout variance against a calibrated two-parameter logistic (2PL) Item Response Theory (IRT) baseline. On a matched subset of 110,000 targets per model, MC-Dropout variance yields an AUC lift of 2.0 to 2.4 percentage points at 80% coverage, while the IRT baseline provides a lift of only 0.41 to 0.46 percentage points—roughly five times smaller. This demonstrates that MC-Dropout is not merely a complex re-implementation of IRT-style confidence but rather uncovers genuine epistemic information that a calibrated psychometric model fundamentally misses.

3. Variance decomposition of epistemic uncertainty (BALD). The authors regress the model's epistemic uncertainty (BALD: total entropy minus expected aleatoric entropy across MC samples) on a nested set of classical psychometric predictors: question difficulty, student ability, IRT-style outcome ambiguity, and curriculum coverage at four granularities (subject, topic, subtopic, construct). Under linear modeling, the entire stack explains at most 3.8% of the variance (2.4% for AKT, 3.8% for SAKT, 1.1% for DKT). Even with a non-linear regressor (5-fold cross-validated random forest), explained variance rises to only 9.8% (SAKT), 10.5% (AKT), and 23.2% (DKT), leaving 76.8% to 90.2% unexplained. The authors conclude this residual is "architecture-specific epistemic content: the model's uncertainty over its own learned representations of a student's unique sequential trajectory, which simpler models cannot capture even with the freedom of non-linear modeling."

4. Responsible KT deployment must rely on the model's intrinsic uncertainty signal. The authors argue that cheaper heuristic alternatives such as population statistics, IRT baselines, or curriculum coverage miss the vast majority of this per-prediction reliability signal by construction.

The paper also provides a concrete deployment example: "an Eedi-style next-question-recommendation pipeline could use the variance signal as a gate. Low-variance predictions would feed directly into the recommender's ranking. High-variance predictions would trigger a fallback (such as a teacher review queue, an easier diagnostic question, or a mastery quiz on the relevant construct), rather than committing to a low-confidence next-question pick."

Limitations acknowledged by the authors include: (1) the headline lifts are improvements on validation-set predictive metrics, not on student learning outcomes or teacher cognitive load; (2) the analysis uses a single dataset (Eedi); (3) the IRT comparison uses a 2-parameter logistic model fit on 30 questions per validation student, and alternative baselines (temperature-scaled softmax, deep ensembles) are not evaluated; (4) the IRT baseline does not address cold-start for students with fewer than 30 prior responses.

The paper concludes that "Selective prediction with model-native epistemic uncertainty is a necessary component of responsible KT deployment. It is complementary to subgroup fairness audits, calibration analysis, and downstream classroom evaluation rather than a replacement for them."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:


Implementation:

I will wrap any trained KT model (DKT, SAKT, AKT, or similar) with a Monte Carlo Dropout (MC-Dropout) inference module. This requires no retraining—I simply keep dropout active during inference and run M=100 stochastic forward passes per prediction. I compute the variance (or standard deviation) of the predicted probabilities across these passes as the uncertainty signal.

Improved capability:

The system can now abstain from making predictions on the 20% most uncertain cases, deferring them to a human teacher. This yields:

  • Accuracy lift: +2.3 to +3.0 percentage points

  • AUC lift: +1.9 to +2.4 percentage points

  • F1 lift: +1.4 to +4.3 percentage points

Crucially, this abstention is targeted: the deferred set has 1.45–1.60× higher error rate than the kept set, meaning the system is not just randomly dropping cases but precisely identifying its own failure modes.

The improved system can:

  • Predict with confidence awareness: It knows when it is uncertain and abstains, improving all key metrics without retraining.

  • Defer intelligently: It routes the most uncertain cases to human teachers, with error rates 1.45–1.60× higher in the deferred set, proving the abstention is targeted.

  • Remain fair: It does not over-defer low-ability students or specific question-difficulty strata, making it safe for real classrooms.

  • Outperform psychometric baselines: It provides 5× the AUC lift of a calibrated 2PL IRT model as a selective-prediction signal, because it captures architecture-specific epistemic content that simpler proxies cannot.

  • Deploy safely: It can be integrated as a gate in recommendation pipelines, ensuring low-confidence predictions never reach students without a human or pedagogical fallback.

Abstract

Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment requires models to know when to defer uncertain predictions to a human teacher. We introduce an intrinsic selective prediction layer for existing KT models using Monte Carlo Dropout (MC-Dropout) to quantify uncertainty. We evaluate this approach across three architectures (DKT, SAKT, and AKT) using the Eedi mathematics dataset. Abstaining on the 20% most uncertain predictions lifts accuracy by 2.3 to 3.0 percentage points, AUC by 1.9 to 2.4 percentage points and F1 by 1.4 to 4.3 percentage points without any retraining. This abstention strategy is highly targeted: the deferred set exhibits 1.45 to 1.60 times the error rate of the kept set. Furthermore, this targeting holds within every question-difficulty quartile and remains fair across student-ability levels. Importantly, MC-Dropout variance gives roughly five times the AUC lift of a calibrated two-parameter logistic (2PL) Item Response Theory (IRT) baseline as a selective-prediction signal. A variance decomposition of the model's epistemic uncertainty (BALD) reveals that the entire classical psychometric stack, comprising question difficulty, student ability, IRT-style outcome ambiguity, and historical curriculum coverage, explains less than 4% of the signal under linear modeling and at most 23% even with a non-linear regressor. This leaves 77% to 90% as architecture-specific epistemic content that MC-Dropout surfaces and simpler proxies cannot recover. Selective prediction with model-native epistemic uncertainty is therefore a necessary component of responsible KT deployment, complementary to subgroup-fairness audits and downstream classroom evaluation rather than a substitute for them.

Sources

Related papers