Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs
summary
In short
The episode examines a paper demonstrating that specialized Knowledge Tracing (KT) models outperform large Language Models (LLMs). These smaller, task-specific models achieve higher accuracy and are significantly faster and cheaper to run than general-purpose LLMs. The conclusion is that specialization is superior for specific tasks like predicting student performance.
Key concepts
- Knowledge Tracing Models
- These are specialized models built to predict how a student will perform on future questions based on their past answers. They are designed for a narrow task—predictive modeling of learning progress—and excel at tracking the student's journey.
- Large Language Models (LLMs)
- LLMs are general-purpose models designed to reason and understand language broadly. While powerful, they struggle with the temporal understanding required for specific tasks like tracking a student's learning journey compared to specialized KT models.
- LLM Knowledge Tracing (LLM KT) Model
- This is a hybrid architecture where an embedding model converts question content into fixed vectors once, and a small transformer handles the temporal prediction. This provides semantic richness without the high runtime cost of full LLM inference.
Terminology used across episodes
This episode discusses
- Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs · Paper Radio
- The Llama 3 Herd of Models · Paper Radio
- MathPrompter: Mathematical Reasoning using Large Language Models
- GPT-4 Technical Report
- Qwen2.5 Technical Report
The paper
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs · Read on arXiv
Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud, Simon Woodhead
Eedi · Learning Engineering Virtual Institute
Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions. One of the key approaches to do this has been through the use of knowledge tracing (KT) models. These are small, domain-specific, temporal models trained on student question-response data. KT models are optimised for high accuracy on specific educational domains and have fast inference and scalable deployments. The rise of Large Language Models (LLMs) motivates us to ask the following questions: (1) How well can LLMs perform at predicting students' future responses to questions? (2) Are LLMs scalable for this domain? (3) How do LLMs compare to KT models on this domain-specific task? In this paper, we compare multiple LLMs and KT models across predictive performance, deployment cost, and inference speed to answer the above questions. We show that KT models outperform LLMs with respect to accuracy and F1 scores on this domain-specific task. Further, we demonstrate that LLMs are orders of magnitude slower than KT models and cost orders of magnitude more to deploy. This highlights the importance of domain-specific models for education prediction tasks and the fact that current closed source LLMs should not be used as a universal solution for all tasks.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs".
Jane: The paper was written by Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud and Simon Woodhead from Eedi and Learning Engineering Virtual Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone! Today we’re looking at a paper that’s going to ruffle some feathers in the AI world. It’s called “Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs.” Jane, I have to say, that title is basically a dare to every big tech company out there.
Jane: It really is, Tom. And I love it. The paper is from researchers at Eedi, which is an online learning platform, and they’ve basically pitted the big general-purpose models like GPT-4o-mini and Gemini against these tiny, specialized models built just for one task: predicting how a student will do on a future math question.
Tom: Right, and the title tells you exactly how it ends. But the scale of the win is what got me. I mean, we’re talking about models with under a million parameters beating models with billions of parameters. That’s like a bicycle beating a Formula One car in a city race.
Jane: That’s a great analogy, because the bicycle is built for the city streets, right? The LLMs are built for everything, but these knowledge tracing models are built for one specific road. They’re looking at a student’s history of answers and trying to figure out if they’ll get the next one right.
Tom: And they’re not just winning by a little. The best specialized model hit about seventy-two point eight percent accuracy, while the best LLM they tested, Gemini, only got to sixty-six point five percent. And some of the LLMs were actually worse than just guessing the average score on every question. That’s a brutal baseline to miss.
Jane: It’s a really important baseline, though. If a model can’t beat "always guess the most common answer," it’s not learning anything about the student. And a few of these huge LLMs couldn’t do it. They just don’t have the temporal understanding—they don’t track the student’s learning journey the way a specialized model is forced to.
Tom: So, the title is basically a spoiler, but the details are in the numbers. And I have to say, as someone who loves efficiency, seeing a zero point seven three million parameter model beat a seven billion parameter model is deeply satisfying.
Jane: It’s the whole argument for specialization in a nutshell. And it sets up the big question we’re going to dig into: if they’re more accurate, are they also actually practical to run? Because that’s where I think the real story is.
Tom: Oh, absolutely. The accuracy gap is one thing, but the cost and speed gap is where this paper really gets wild. Stick around, because we’re about to talk about latency and what it actually costs to serve a hundred thousand students.
Summary: Tom: So, Jane, we’ve established that the specialized models win on accuracy. But the paper’s summary really hammers home the other two words in the title: faster and cheaper. And honestly, the numbers there are almost comical.
Jane: They really are. Let’s talk about speed first. The knowledge tracing models—DKT, SAKT, and their new one, LLM KT—they all predict a student’s next answer in under a quarter of a second. That’s per student, for their whole history. Meanwhile, GPT-4o-mini takes about three seconds, which is slow but manageable.
Tom: But then you get to the open-source models they hosted themselves. Qwen2 point 5-7B took over three thousand two hundred seconds per student. That’s almost an hour. For one student’s prediction. You can’t run an adaptive learning platform if the model takes an hour to decide what question to show next.
Jane: Exactly. And that’s where the "cheaper" part comes in. They calculated the annual cost to serve one hundred thousand students, each getting forty predictions. The specialized models cost less than two dollars a year. Two dollars, Tom.
Tom: And the LLMs? GPT-4o-mini came in around,300 a year. But the self-hosted ones were brutal. Qwen2 point 5-7B would cost almost,000 a year. Llama-1B, even the fine-tuned one, was nearly,000. So we’re talking about a six hundred to twelve thousand times cost difference.
Jane: And here’s the kicker—they’re paying more for worse results. The fine-tuned Llama-1B got seventy-one percent accuracy, which is actually decent, but it costs,000 and takes half an hour per student. The specialized SAKT model got seventy-two point seven percent accuracy, costs under a dollar, and runs in a tenth of a second.
Tom: It’s not even a trade-off. It’s just a strictly better option. And I think that’s the point the authors are making. They’re not saying LLMs are useless—they’re saying that for this specific task, using a general-purpose model is like using a sledgehammer to hang a picture frame.
Jane: Right. And the sledgehammer is also on fire and costs a fortune. But I do want to bring in Lu and Meng here, because I think they’ll have different takes on why this gap exists and whether it can be closed. Lu, you work with these models every day—does this surprise you?
Lu: Honestly, Jane, it doesn’t surprise me at all. The LLMs are solving a much harder problem. They’re trying to be a general reasoner. The KT models are solving a narrow temporal prediction problem, and they can encode the structure of that problem directly into the architecture. The LLM has to rediscover that structure from text prompts, which is incredibly inefficient.
Meng: And from a deployment standpoint, it’s not just the inference cost. It’s the memory footprint, the GPU requirements, the cold-start latency. Running a 7B model on CPU is painful. The specialized models are so small you can run them on basically anything. That’s a massive operational advantage.
Tom: Great points, both of you. So we’ve got accuracy, speed, and cost all favoring the specialists. But the paper doesn’t stop at just comparing. They also built their own model, and that’s where things get interesting. Let’s talk about that next.
Improvements: Tom: So, Jane, the paper doesn’t just say "our old models are better." They actually built a new model, and it’s the best of both worlds. It’s called LLM KT, and it’s a clever hybrid that uses LLM embeddings but keeps the temporal modeling in a tiny transformer.
Jane: Right, and this is the part I found really smart. Instead of feeding the LLM the whole student history and asking it to reason, they use a small embedding model—Qwen three 0 point 6B—to convert the question text, the construct, the misconceptions, all of that into fixed vectors. They cache those embeddings offline.
Tom: So the LLM does the heavy lifting of understanding the question content, but it happens once, ahead of time. Then at runtime, the tiny KT model just looks up those pre-computed vectors and does its temporal thing. No LLM inference in the loop at all.
Jane: Exactly. And the result is that LLM KT gets the best accuracy in the whole paper—seventy-two point eight percent, just barely edging out SAKT at seventy-two point seven percent. But it also keeps that sub-zero point two five second latency and the sub- annual cost. They got the semantic richness of an LLM without paying the runtime price.
Meng: That’s a really elegant engineering solution. You’re essentially using the LLM as a feature extractor, not as a predictor. And because the embeddings are cached, you could even swap out the embedding model later if a better one comes along, without retraining the temporal model.
Lu: And I think this points to the future. We’re going to see more of this pattern—using LLMs to generate rich representations, but then using small, specialized models to do the actual prediction. It’s a division of labor. The LLM handles language understanding; the specialized model handles the student’s learning state.
Tom: It also solves a practical problem they mentioned in the paper. The zero-shot Llama-1B model kept failing to follow the prompt format. It would output random text instead of just "Yes" or "No." By taking the language generation out of the loop entirely, you eliminate that failure mode completely.
Jane: That’s a great point. The fine-tuned Llama did better—seventy-one percent accuracy—but it still required prompt engineering and careful output parsing. The LLM KT model just sidesteps all of that. It’s a much more robust system for production.
Meng: And the cost comparison really shows it. The fine-tuned Llama cost,000 a year to serve. The LLM KT model costs.73. For the same scale of deployment, you’re getting better accuracy for a fraction of the cost. That’s the kind of improvement that actually makes a product viable.
Tom: So the improvement isn’t just a better model—it’s a better architecture for the problem. And it makes me wonder, Lu, where does this leave the big LLMs? Are they just going to be relegated to embedding generators?
Lu: Not relegated, but specialized. They’re still the best at understanding language. But for prediction tasks with temporal structure, we’re going to see more and more of these hybrid systems. The LLM provides the semantic understanding; the small model provides the temporal reasoning. It’s a partnership, not a competition.
Jane: And that’s a much more nuanced takeaway than "LLMs are bad." They’re not bad—they’re just not the right tool for every job. And this paper proves that with numbers, not just opinions. Let’s bring in Lalam to get a bigger picture on what this means for education technology.
Conclusion: Tom: Alright, we’ve covered the accuracy, the speed, the cost, and the clever hybrid architecture. Let’s wrap this up. Jane, what’s the big picture takeaway from "Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs"?
Jane: The big picture is that specialization wins for specific tasks. The paper shows that if you want to predict student performance, you should use a model built for that. The LLMs, even the best ones, just can’t match the temporal modeling that KT models do natively.
Tom: And the practical implications are huge. For an EdTech platform trying to serve millions of students, the cost difference between and,000 a year is the difference between a viable product and a money pit. It’s not just about accuracy—it’s about whether you can actually deploy it.
Meng: And from an engineering standpoint, the LLM KT model is the kind of solution I love. It uses the best of both worlds—LLM embeddings for semantic understanding, small transformer for temporal prediction—and it does it in a way that’s actually deployable on commodity hardware.
Lu: I’d add that this is a template for other domains too. Anywhere you have temporal sequences with rich textual context—health records, financial transactions, even recommendation systems—this hybrid pattern could apply. Use the LLM to understand the content, use a small model to track the state.
Lalam: And I think the cultural impact is about access. When models are this cheap and fast, they can be deployed in resource-constrained environments—schools in developing regions, rural areas with limited connectivity. The barrier to personalized learning drops dramatically. That’s the real win.
Tom: That’s a beautiful way to put it, Lalam. The paper isn’t just about winning a benchmark—it’s about making adaptive learning accessible to everyone. And that’s a mission we can all get behind.
Jane: Absolutely. So, to sum up: specialized KT models beat LLMs on accuracy, speed, and cost. The new LLM KT model shows you can have the best of both worlds by using LLM embeddings offline. And the implications for education are profound.
Tom: And with that, we’re going to say goodbye to this paper. It’s been a great discussion, and I think we’ve all learned something about the value of specialization. Thanks to Lu, Meng, and Lalam for joining us.
Jane: And thanks to all our listeners. We’ll be back soon with another paper, and we’ll see if it can live up to the standard this one set. Until then, keep learning, keep questioning, and remember—sometimes the smallest model is the smartest choice.
Tom: See you next time, everyone!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language