Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
summary
The gist
This paper investigates whether large language models (LLMs) can effectively model "incorrect yet plausible reasoning" through the task of multiple-choice distractor generation.
In short
Researchers investigated whether LLMs can mimic human mistakes by generating plausible multiple-choice distractors. Testing models like DeepSeek-V3 and GLM-4 on the Eedi math dataset, the study found they use a 'solve-first' strategy. Providing the correct answer in prompts improved performance by 8%, helping models anchor their error simulations.
Key concepts
- Distractors
- In multiple-choice tests, distractors are incorrect options designed to be plausible enough that a confused student might actually pick them. The research explores whether AI can generate these realistic wrong answers to better simulate human cognitive patterns and mistakes.
- Solve-first strategy
- This is a method where an AI model solves a math problem correctly in its own reasoning process before intentionally taking a wrong turn. By establishing the correct path first, the model can more reliably simulate specific human errors.
- Error simulation
- This involves using structured reasoning to map out the anatomy of a mistake. Instead of stumbling onto wrong answers by accident, the model uses logic—such as simulating a student who forgets to divide—to provide transparency in its reasoning process.
Terminology used across episodes
This episode discusses
- Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation · Paper Radio
- Math Multiple Choice Question Generation via Human-Large Language Model Collaboration
- Novice Learner and Expert Tutor: Evaluating Math Reasoning Abilities of Large Language Models with Misconceptions
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
The paper
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation · Read on arXiv
ETH Zürich · Bocconi University · University of Central Florida
Modeling student misconceptions in a realistic manner is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating distractor answers for multiple-choice questions (MCQs), a task that requires producing answers that are incorrect, yet plausible. We introduce a taxonomy over reasoning strategies for distractor generation that is grounded in learning-science literature and empirical observation, which we apply to LLM-generated reasoning traces across math and science MCQs. On the math dataset, we find that models follow a misconception-based process with potentially high diagnostic value: they recover the correct solution, articulate student errors, simulate them, and select plausible candidates. On the science dataset, on the other hand, they tend to follow a less robust approach based on semantic similarity to the correct answer. We find the most frequent failure modes to be that the model is unable to generate a correct solution or that it discards plausible distractor candidates when performing selection. Providing the correct solution in the prompt yields a relative improvement of 6.4% in alignment with human-authored distractors, highlighting the critical role of anchoring distractor generation to the correct solution. Together, our findings offer an interpretable view of how LLMs model incorrect student reasoning.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation".
Jane: The paper was written by Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar et al. from ETH Zürich and Bocconi University and University of Central Florida.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We are starting today with a fascinating new paper titled "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation."
Jane: It's a really interesting study, Tom, because it looks at how researchers from ETH Zürich and other institutions are testing if AI can actually mimic human mistakes.
Tom: They focus specifically on those tricky wrong answers in multiple-choice tests.
Jane: Those are called distractors, and they have to be plausible enough that a confused student might actually pick them.
Lu: This goes way beyond just making better tests for schools.
Lu: If we can teach these models to simulate the specific ways a human mind trips up, we're building AI that understands our actual cognitive patterns.
Meng: I wonder about the practical side of that, Lu.
Meng: Does this mean we can eventually automate the creation of massive, high-quality practice exams for students everywhere?
Lu: It could lead to much more personalized learning than anything we've seen before.
Meng: If it works, it would definitely take a huge load off teachers who spend hours writing these questions.
Lalam: There is a deeper shift happening here as well.
Lalam: By teaching machines to model errors, we move our culture away from demanding instant perfection and toward valuing the process of learning through mistakes.
Jane: That makes the technology feel much more like a supportive tutor than just a machine that marks things wrong.
Tom: It's a great starting point, Jane, and it leads us right into how these models actually think during the process.
Summary: Tom: Building on that idea of how they think, the researchers tested DeepSeek-V3 point 2 and GLM-four point seven using a math dataset called Eedi.
Jane: They wanted to see if these models were just guessing or if they had a specific method for creating those distractors.
Tom: It turns out they follow what the paper calls a "solve-first" strategy.
Jane: Basically, the model solves the math problem correctly in its own reasoning process before it even tries to be wrong.
Lu: That's such a clever way for a model to operate!
Lu: It is like the model builds a perfect map of the correct path and then intentionally takes a wrong turn at a specific point.
Meng: I was looking at the taxonomy they used to track this behavior.
Meng: They broke it down into steps like "error description" and "error simulation" to see exactly how the model's logic unfolds.
Lu: It really is like they are mapping out the anatomy of a mistake.
Meng: Having that kind of structured reasoning makes the whole process much more reliable for an engineer to implement.
Jane: It shows the model isn't just stumbling onto wrong answers by accident.
Lalam: That structure provides a huge amount of transparency for us.
Lalam: When a model says it is simulating a student who forgets to divide, we can actually see the logic behind its mistake.
Tom: Seeing that reasoning process is what helps us understand if the AI is actually engaging with human logic.
Improvements: Tom: Even with that structured approach, the paper shows there are still some clear places where these models fail.
Jane: One major problem they identified was when the model struggled to recover the correct solution in the first place.
Tom: If it can't get the right answer, it has nothing solid to deviate from.
Jane: It is like trying to plan a detour when you don't even know where your original destination was.
Lu: They actually found a very direct way to help with this issue.
Lu: If you simply provide the correct answer in the prompt, the model's performance improves by eight percent.
Meng: That is a very practical suggestion for anyone building these tools.
Meng: You can use the correct answer as an anchor so the model stays grounded while it explores errors.
Lu: We could eventually use even more advanced anchors to guide much more complex simulations.
Meng: I do worry about the cost of all those extra reasoning tokens, though.
Meng: Generating all that extra text for thinking might make the system slower or more expensive to run in a real app.
Lalam: The trade-off might be worth it for the clarity it provides.
Lalam: That extra explanation bridges the gap between cold machine logic and our own human intuition.
Jane: It makes the whole experience feel much more natural for a student using it.
Tom: We'll see if those improvements can bridge that gap in the next stage of development.
Conclusion: Tom: We have covered a lot today, from how models use error injection to the ways we can improve their accuracy.
Jane: It really highlights how much progress we've made in moving toward true cognitive modeling.
Tom: Before we go, I want to hear one last thought from the team on "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation."
Lu: This research is a massive step toward a world where AI understands the beauty of human trial and error.
Meng: I am excited to see how we can take these structured reasoning traces and turn them into stable, production-ready educational tools.
Lalam: My hope is that this leads to a culture where every learner feels understood because the technology recognizes their unique way of thinking.
Tom: That is a perfect place to stop, Jane.
Jane: I agree, Tom! Thanks for joining us everyone, and we will see you next time!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language