Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs

arXiv:2603.02830 · cs.CL, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs".

Jane: The paper was written by Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud and Simon Woodhead from Eedi and Learning Engineering Virtual Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone! Today we’re looking at a paper that’s going to ruffle some feathers in the AI world. It’s called “Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs.” Jane, I have to say, that title is basically a dare to every big tech company out there.

Jane: It really is, Tom. And I love it. The paper is from researchers at Eedi, which is an online learning platform, and they’ve basically pitted the big general-purpose models like GPT-4o-mini and Gemini against these tiny, specialized models built just for one task: predicting how a student will do on a future math question.

Tom: Right, and the title tells you exactly how it ends. But the scale of the win is what got me. I mean, we’re talking about models with under a million parameters beating models with billions of parameters. That’s like a bicycle beating a Formula One car in a city race.

Jane: That’s a great analogy, because the bicycle is built for the city streets, right? The LLMs are built for everything, but these knowledge tracing models are built for one specific road. They’re looking at a student’s history of answers and trying to figure out if they’ll get the next one right.

Tom: And they’re not just winning by a little. The best specialized model hit about seventy-two point eight percent accuracy, while the best LLM they tested, Gemini, only got to sixty-six point five percent. And some of the LLMs were actually worse than just guessing the average score on every question. That’s a brutal baseline to miss.

Jane: It’s a really important baseline, though. If a model can’t beat "always guess the most common answer," it’s not learning anything about the student. And a few of these huge LLMs couldn’t do it. They just don’t have the temporal understanding—they don’t track the student’s learning journey the way a specialized model is forced to.

Tom: So, the title is basically a spoiler, but the details are in the numbers. And I have to say, as someone who loves efficiency, seeing a zero point seven three million parameter model beat a seven billion parameter model is deeply satisfying.

Jane: It’s the whole argument for specialization in a nutshell. And it sets up the big question we’re going to dig into: if they’re more accurate, are they also actually practical to run? Because that’s where I think the real story is.

Tom: Oh, absolutely. The accuracy gap is one thing, but the cost and speed gap is where this paper really gets wild. Stick around, because we’re about to talk about latency and what it actually costs to serve a hundred thousand students.

Summary: Tom: So, Jane, we’ve established that the specialized models win on accuracy. But the paper’s summary really hammers home the other two words in the title: faster and cheaper. And honestly, the numbers there are almost comical.

Jane: They really are. Let’s talk about speed first. The knowledge tracing models—DKT, SAKT, and their new one, LLM KT—they all predict a student’s next answer in under a quarter of a second. That’s per student, for their whole history. Meanwhile, GPT-4o-mini takes about three seconds, which is slow but manageable.

Tom: But then you get to the open-source models they hosted themselves. Qwen2 point 5-7B took over three thousand two hundred seconds per student. That’s almost an hour. For one student’s prediction. You can’t run an adaptive learning platform if the model takes an hour to decide what question to show next.

Jane: Exactly. And that’s where the "cheaper" part comes in. They calculated the annual cost to serve one hundred thousand students, each getting forty predictions. The specialized models cost less than two dollars a year. Two dollars, Tom.

Tom: And the LLMs? GPT-4o-mini came in around,300 a year. But the self-hosted ones were brutal. Qwen2 point 5-7B would cost almost,000 a year. Llama-1B, even the fine-tuned one, was nearly,000. So we’re talking about a six hundred to twelve thousand times cost difference.

Jane: And here’s the kicker—they’re paying more for worse results. The fine-tuned Llama-1B got seventy-one percent accuracy, which is actually decent, but it costs,000 and takes half an hour per student. The specialized SAKT model got seventy-two point seven percent accuracy, costs under a dollar, and runs in a tenth of a second.

Tom: It’s not even a trade-off. It’s just a strictly better option. And I think that’s the point the authors are making. They’re not saying LLMs are useless—they’re saying that for this specific task, using a general-purpose model is like using a sledgehammer to hang a picture frame.

Jane: Right. And the sledgehammer is also on fire and costs a fortune. But I do want to bring in Lu and Meng here, because I think they’ll have different takes on why this gap exists and whether it can be closed. Lu, you work with these models every day—does this surprise you?

Lu: Honestly, Jane, it doesn’t surprise me at all. The LLMs are solving a much harder problem. They’re trying to be a general reasoner. The KT models are solving a narrow temporal prediction problem, and they can encode the structure of that problem directly into the architecture. The LLM has to rediscover that structure from text prompts, which is incredibly inefficient.

Meng: And from a deployment standpoint, it’s not just the inference cost. It’s the memory footprint, the GPU requirements, the cold-start latency. Running a 7B model on CPU is painful. The specialized models are so small you can run them on basically anything. That’s a massive operational advantage.

Tom: Great points, both of you. So we’ve got accuracy, speed, and cost all favoring the specialists. But the paper doesn’t stop at just comparing. They also built their own model, and that’s where things get interesting. Let’s talk about that next.

Improvements: Tom: So, Jane, the paper doesn’t just say "our old models are better." They actually built a new model, and it’s the best of both worlds. It’s called LLM KT, and it’s a clever hybrid that uses LLM embeddings but keeps the temporal modeling in a tiny transformer.

Jane: Right, and this is the part I found really smart. Instead of feeding the LLM the whole student history and asking it to reason, they use a small embedding model—Qwen three 0 point 6B—to convert the question text, the construct, the misconceptions, all of that into fixed vectors. They cache those embeddings offline.

Tom: So the LLM does the heavy lifting of understanding the question content, but it happens once, ahead of time. Then at runtime, the tiny KT model just looks up those pre-computed vectors and does its temporal thing. No LLM inference in the loop at all.

Jane: Exactly. And the result is that LLM KT gets the best accuracy in the whole paper—seventy-two point eight percent, just barely edging out SAKT at seventy-two point seven percent. But it also keeps that sub-zero point two five second latency and the sub- annual cost. They got the semantic richness of an LLM without paying the runtime price.

Meng: That’s a really elegant engineering solution. You’re essentially using the LLM as a feature extractor, not as a predictor. And because the embeddings are cached, you could even swap out the embedding model later if a better one comes along, without retraining the temporal model.

Lu: And I think this points to the future. We’re going to see more of this pattern—using LLMs to generate rich representations, but then using small, specialized models to do the actual prediction. It’s a division of labor. The LLM handles language understanding; the specialized model handles the student’s learning state.

Tom: It also solves a practical problem they mentioned in the paper. The zero-shot Llama-1B model kept failing to follow the prompt format. It would output random text instead of just "Yes" or "No." By taking the language generation out of the loop entirely, you eliminate that failure mode completely.

Jane: That’s a great point. The fine-tuned Llama did better—seventy-one percent accuracy—but it still required prompt engineering and careful output parsing. The LLM KT model just sidesteps all of that. It’s a much more robust system for production.

Meng: And the cost comparison really shows it. The fine-tuned Llama cost,000 a year to serve. The LLM KT model costs.73. For the same scale of deployment, you’re getting better accuracy for a fraction of the cost. That’s the kind of improvement that actually makes a product viable.

Tom: So the improvement isn’t just a better model—it’s a better architecture for the problem. And it makes me wonder, Lu, where does this leave the big LLMs? Are they just going to be relegated to embedding generators?

Lu: Not relegated, but specialized. They’re still the best at understanding language. But for prediction tasks with temporal structure, we’re going to see more and more of these hybrid systems. The LLM provides the semantic understanding; the small model provides the temporal reasoning. It’s a partnership, not a competition.

Jane: And that’s a much more nuanced takeaway than "LLMs are bad." They’re not bad—they’re just not the right tool for every job. And this paper proves that with numbers, not just opinions. Let’s bring in Lalam to get a bigger picture on what this means for education technology.

Conclusion: Tom: Alright, we’ve covered the accuracy, the speed, the cost, and the clever hybrid architecture. Let’s wrap this up. Jane, what’s the big picture takeaway from "Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs"?

Jane: The big picture is that specialization wins for specific tasks. The paper shows that if you want to predict student performance, you should use a model built for that. The LLMs, even the best ones, just can’t match the temporal modeling that KT models do natively.

Tom: And the practical implications are huge. For an EdTech platform trying to serve millions of students, the cost difference between and,000 a year is the difference between a viable product and a money pit. It’s not just about accuracy—it’s about whether you can actually deploy it.

Meng: And from an engineering standpoint, the LLM KT model is the kind of solution I love. It uses the best of both worlds—LLM embeddings for semantic understanding, small transformer for temporal prediction—and it does it in a way that’s actually deployable on commodity hardware.

Lu: I’d add that this is a template for other domains too. Anywhere you have temporal sequences with rich textual context—health records, financial transactions, even recommendation systems—this hybrid pattern could apply. Use the LLM to understand the content, use a small model to track the state.

Lalam: And I think the cultural impact is about access. When models are this cheap and fast, they can be deployed in resource-constrained environments—schools in developing regions, rural areas with limited connectivity. The barrier to personalized learning drops dramatically. That’s the real win.

Tom: That’s a beautiful way to put it, Lalam. The paper isn’t just about winning a benchmark—it’s about making adaptive learning accessible to everyone. And that’s a mission we can all get behind.

Jane: Absolutely. So, to sum up: specialized KT models beat LLMs on accuracy, speed, and cost. The new LLM KT model shows you can have the best of both worlds by using LLM embeddings offline. And the implications for education are profound.

Tom: And with that, we’re going to say goodbye to this paper. It’s been a great discussion, and I think we’ve all learned something about the value of specialization. Thanks to Lu, Meng, and Lalam for joining us.

Jane: And thanks to all our listeners. We’ll be back soon with another paper, and we’ll see if it can live up to the standard this one set. Until then, keep learning, keep questioning, and remember—sometimes the smallest model is the smartest choice.

Tom: See you next time, everyone!

Prarthana Bhattacharyya, Joshua Mitton, Ralph Abboud, Simon Woodhead

Eedi · Learning Engineering Virtual Institute

cs.CL, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 7 pages, 6 figures. Prarthana Bhattacharyya and Joshua Mitton contributed equally to this work

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

Key concepts

Knowledge Tracing Models
These are specialized models built to predict how a student will perform on future questions based on their past answers. They are designed for a narrow task—predictive modeling of learning progress—and excel at tracking the student's journey.
Large Language Models (LLMs)
LLMs are general-purpose models designed to reason and understand language broadly. While powerful, they struggle with the temporal understanding required for specific tasks like tracking a student's learning journey compared to specialized KT models.
LLM Knowledge Tracing (LLM KT) Model
This is a hybrid architecture where an embedding model converts question content into fixed vectors once, and a small transformer handles the temporal prediction. This provides semantic richness without the high runtime cost of full LLM inference.

Terminology

Summary

Summary

This paper presents a systematic comparison between specialised Knowledge Tracing (KT) models and general-purpose Large Language Models (LLMs) for the task of predicting students' future responses to questions. The authors address three research questions: (1) How well can LLMs predict students' future responses? (2) Are LLMs scalable for this domain in terms of latency and cost? (3) How do LLMs compare to specialised KT models on this task?

The task is framed as a binary classification problem: given a student's response history to past questions, the model must predict whether the student will answer a future question correctly or incorrectly. The models have access to question text or question id, construct text or construct id, and for each question, misconception text for each possible answer and explanations of the answers. To avoid cold-start issues, predictions are not made on the first 10 questions, and each student requires 40 answer predictions. The evaluation uses a real-world dataset from an online learning platform with 512,000 training responses (12,800 students, 4,252 questions) and 64,000 validation responses (1,600 students, 4,104 questions). The dataset bias (percentage of correct responses) is 65.8% for training and 66.5% for validation. Students in training and validation sets are different, testing generalisability to new students.

The models compared include: DKT (Deep Knowledge Tracing), SAKT (Self-Attentive Knowledge Tracing), LLM KT (a custom-built encoder-decoder temporal transformer using Qwen 3 0.6B embeddings as fixed feature extractors, with no LLM involved at inference time), GPT-4o-mini, Gemini-2.5-flash-lite, Qwen2.5-7B-Instruct, Llama-1B zero-shot, and Llama-1B with LoRA fine-tuning. For LLMs, a constrained prompt design is used: "The student answered the following questions: Question ID [id] ([question text with choices A/B/C/D]), with construct ID [id] ([construct description]): answered [correctly/incorrectly]... Predict whether the student will answer [new question with full context] correctly or not. Answer ONLY with the word 'Yes' or the word 'No'." Metrics used are accuracy and macro-averaged F1 score.

Results show that specialised KT models outperform LLMs on predictive performance. LLM KT achieves the highest accuracy at 72.8%, narrowly outperforming SAKT (72.7%) and DKT (71.8%). All domain-specific KT models outperform the general-purpose LLMs: GPT-4o-mini (58.6%), Qwen2.5-7B-Instruct (64.6%), Gemini-2.5-flash-lite (66.5%), Llama-1B LoRA fine-tuned (71.0%), and Llama-1B zero-shot (33.5%). For context, a dataset bias baseline (predicting the average correctness rate for each question) achieves 66.5% accuracy, and several LLMs fail to surpass this naive baseline. For F1 score, LLM KT leads with 0.674, followed by SAKT (0.669) and DKT (0.650). LLMs score lower: GPT-4o-mini (0.579), Qwen2.5-7B-Instruct (0.533), Gemini-2.5-flash-lite (0.527), Llama-1B LoRA fine-tune (0.592), and Llama-1B zero-shot (0.251). The authors note that Llama-1B zero-shot frequently fails to follow the prompt format, producing responses other than the expected 'correct' or 'incorrect,' which are treated as incorrect predictions. The remaining LLMs adhere to the prompt format but struggle to predict student responses accurately, failing to infer that students tend to get more answers correct in general.

Regarding latency, KT models deliver fast predictions with latencies under 0.25 seconds per student (DKT: 0.09s, SAKT: 0.10s, LLM KT: 0.23s). In contrast, LLMs are orders of magnitude slower: GPT-4o-mini takes 3.1 seconds, Gemini-2.5-flash-lite takes 128 seconds, Qwen2.5-7B-Instruct requires 3,299 seconds per student, and Llama-1B variants take 1,598.8 seconds per student. Model sizes also differ dramatically: KT models range from 0.58M to 0.85M parameters (LLM KT has 0.73M), while LLMs have 1B to 8B parameters (GPT-4o-mini: 8B, Gemini-2.5-flash-lite: 4B, Qwen2.5-7B-Instruct: 7B, Llama-1B variants: 1B).

For cost analysis, the authors calculate annual inference cost for 100,000 students, each receiving 40 predictions per year. Specialised KT models cost less than 2 per year (DKT: 0.675, SAKT: 0.75, LLM KT: 1.73). In contrast, GPT-4o-mini costs approximately 2,322 per year, Gemini-2.5-flash-lite costs 1,230 per year, Llama-1B variants cost 11,991 per year, and Qwen2.5-7B-Instruct reaches 24,741 per year. This means KT models are 615–12,400 times cheaper than LLMs for the same task, while also offering higher accuracy. For Gemini-2.5-flash-lite, daily rate limits of 10,000 requests per day on Tier-1 restricted testing to 200 students, from which costs were extrapolated.

The authors conclude that despite the rapid rise of general-purpose LLMs, our findings demonstrate that specialised KT models remain the best choice for predicting student responses in EdTech settings. KT models outperform LLMs on accuracy and F1 score, are over 600 times cheaper to deploy at scale, and offer millisecond-level latency without requiring GPUs. The paper highlights the importance of domain-specific models for education prediction tasks and the fact that current closed source LLMs should not be used as a universal solution for all tasks. The authors note that LLMs excel at broad reasoning tasks but fall short when applied to student interaction data, and recommend that LLMs should be employed where they demonstrably reduce costs or improve learning outcomes, not as a default choice.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace pure LLM inference with a two-stage pipeline:

  • Stage 1 (Offline): Use a small embedding model (e.g., Qwen 0.6B) to convert question text, construct text, misconception text, and explanation text into fixed vector representations. Cache these embeddings.

  • Stage 2 (Online): Feed these embeddings into a compact temporal transformer (0.73M parameters) that models student knowledge state over time.

Resulting capability: Achieves 72.8% accuracy and 0.674 F1 score—outperforming all tested LLMs (GPT-4o-mini: 58.6%, Gemini-2.5-flash-lite: 66.5%, Qwen2.5-7B: 64.6%) while running in under 0.25 seconds per student and costing less than 2/year for 100,000 students.

Abstract

Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective interventions. One of the key approaches to do this has been through the use of knowledge tracing (KT) models. These are small, domain-specific, temporal models trained on student question-response data. KT models are optimised for high accuracy on specific educational domains and have fast inference and scalable deployments. The rise of Large Language Models (LLMs) motivates us to ask the following questions: (1) How well can LLMs perform at predicting students' future responses to questions? (2) Are LLMs scalable for this domain? (3) How do LLMs compare to KT models on this domain-specific task? In this paper, we compare multiple LLMs and KT models across predictive performance, deployment cost, and inference speed to answer the above questions. We show that KT models outperform LLMs with respect to accuracy and F1 scores on this domain-specific task. Further, we demonstrate that LLMs are orders of magnitude slower than KT models and cost orders of magnitude more to deploy. This highlights the importance of domain-specific models for education prediction tasks and the fact that current closed source LLMs should not be used as a universal solution for all tasks.

Sources

Related papers