Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation".
Jane: The paper was written by Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar et al. from ETH Zürich and Bocconi University and University of Central Florida.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We are starting today with a fascinating new paper titled "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation."
Jane: It's a really interesting study, Tom, because it looks at how researchers from ETH Zürich and other institutions are testing if AI can actually mimic human mistakes.
Tom: They focus specifically on those tricky wrong answers in multiple-choice tests.
Jane: Those are called distractors, and they have to be plausible enough that a confused student might actually pick them.
Lu: This goes way beyond just making better tests for schools.
Lu: If we can teach these models to simulate the specific ways a human mind trips up, we're building AI that understands our actual cognitive patterns.
Meng: I wonder about the practical side of that, Lu.
Meng: Does this mean we can eventually automate the creation of massive, high-quality practice exams for students everywhere?
Lu: It could lead to much more personalized learning than anything we've seen before.
Meng: If it works, it would definitely take a huge load off teachers who spend hours writing these questions.
Lalam: There is a deeper shift happening here as well.
Lalam: By teaching machines to model errors, we move our culture away from demanding instant perfection and toward valuing the process of learning through mistakes.
Jane: That makes the technology feel much more like a supportive tutor than just a machine that marks things wrong.
Tom: It's a great starting point, Jane, and it leads us right into how these models actually think during the process.
Summary: Tom: Building on that idea of how they think, the researchers tested DeepSeek-V3 point 2 and GLM-four point seven using a math dataset called Eedi.
Jane: They wanted to see if these models were just guessing or if they had a specific method for creating those distractors.
Tom: It turns out they follow what the paper calls a "solve-first" strategy.
Jane: Basically, the model solves the math problem correctly in its own reasoning process before it even tries to be wrong.
Lu: That's such a clever way for a model to operate!
Lu: It is like the model builds a perfect map of the correct path and then intentionally takes a wrong turn at a specific point.
Meng: I was looking at the taxonomy they used to track this behavior.
Meng: They broke it down into steps like "error description" and "error simulation" to see exactly how the model's logic unfolds.
Lu: It really is like they are mapping out the anatomy of a mistake.
Meng: Having that kind of structured reasoning makes the whole process much more reliable for an engineer to implement.
Jane: It shows the model isn't just stumbling onto wrong answers by accident.
Lalam: That structure provides a huge amount of transparency for us.
Lalam: When a model says it is simulating a student who forgets to divide, we can actually see the logic behind its mistake.
Tom: Seeing that reasoning process is what helps us understand if the AI is actually engaging with human logic.
Improvements: Tom: Even with that structured approach, the paper shows there are still some clear places where these models fail.
Jane: One major problem they identified was when the model struggled to recover the correct solution in the first place.
Tom: If it can't get the right answer, it has nothing solid to deviate from.
Jane: It is like trying to plan a detour when you don't even know where your original destination was.
Lu: They actually found a very direct way to help with this issue.
Lu: If you simply provide the correct answer in the prompt, the model's performance improves by eight percent.
Meng: That is a very practical suggestion for anyone building these tools.
Meng: You can use the correct answer as an anchor so the model stays grounded while it explores errors.
Lu: We could eventually use even more advanced anchors to guide much more complex simulations.
Meng: I do worry about the cost of all those extra reasoning tokens, though.
Meng: Generating all that extra text for thinking might make the system slower or more expensive to run in a real app.
Lalam: The trade-off might be worth it for the clarity it provides.
Lalam: That extra explanation bridges the gap between cold machine logic and our own human intuition.
Jane: It makes the whole experience feel much more natural for a student using it.
Tom: We'll see if those improvements can bridge that gap in the next stage of development.
Conclusion: Tom: We have covered a lot today, from how models use error injection to the ways we can improve their accuracy.
Jane: It really highlights how much progress we've made in moving toward true cognitive modeling.
Tom: Before we go, I want to hear one last thought from the team on "Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation."
Lu: This research is a massive step toward a world where AI understands the beauty of human trial and error.
Meng: I am excited to see how we can take these structured reasoning traces and turn them into stable, production-ready educational tools.
Lalam: My hope is that this leads to a culture where every learner feels understood because the technology recognizes their unique way of thinking.
Tom: That is a perfect place to stop, Jane.
Jane: I agree, Tom! Thanks for joining us everyone, and we will see you next time!
ETH Zürich · Bocconi University · University of Central Florida
cs.CL, cs.AI, cs.HC
Submitted: 2026-03-16
Updated: 2026-09-15
Comments: Accepted to the Findings of EMNLP 2026
Code: https://github.com/eth-lre/llm-student-modeling-strategies
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 85/100
The gist: This paper investigates whether large language models (LLMs) can effectively model "incorrect yet plausible reasoning" through the task of multiple-choice distractor generation.
Key concepts
- Distractors
- In multiple-choice tests, distractors are incorrect options designed to be plausible enough that a confused student might actually pick them. The research explores whether AI can generate these realistic wrong answers to better simulate human cognitive patterns and mistakes.
- Solve-first strategy
- This is a method where an AI model solves a math problem correctly in its own reasoning process before intentionally taking a wrong turn. By establishing the correct path first, the model can more reliably simulate specific human errors.
- Error simulation
- This involves using structured reasoning to map out the anatomy of a mistake. Instead of stumbling onto wrong answers by accident, the model uses logic—such as simulating a student who forgets to divide—to provide transparency in its reasoning process.
Terminology
Summary
This paper investigates whether large language models (LLMs) can effectively model incorrect yet plausible reasoning
through the task of multiple-choice distractor generation. Understanding how models simulate student misconceptions is vital for developing AI-driven educational tools, such as diagnostic systems and automated test design, which require a faithful representation of both correct knowledge and systematic errors.
A taxonomy of strategies
To analyze the reasoning procedures used by state-of-the-art LLMs, the authors introduce a broadly applicable framework for analyzing and comparing reasoning traces.
This taxonomy integrates learning-science theory with empirical observations to categorize how models arrive at distractor candidates. The researchers applied this taxonomy to 240 traces from DeepSeek-V3.2 and GLM-4.7 using the Eedi Math MCQ Dataset, identifying several key strategies:
-
INTER: Clarifying task instructions or requirements.
-
CORR: Computing or referencing the correct solution.
-
ERR DESC: Describing common misconceptions or procedural errors at an abstract level.
-
INST: Stating the incorrect outcome resulting from an error.
-
ERR SIM: Deliberately verbalizing an incorrect procedure through error simulation.
-
PLAUS: Evaluating how likely a candidate answer is to be chosen by a student.
-
CURATE: Selecting and refining a set of distinct, incorrect distractors.
-
RECON: Revisiting earlier choices or alternative interpretations.
LLM reasoning patterns
The study examines whether LLM processes align with best practices in the learning sciences.
The results reveal a surprising alignment
between model behavior and recommended misconception-based design practices. Rather than relying on similarity-based heuristics—which modify surface features of the correct answer without modeling underlying reasoning—the models predominantly follow a solve-first, error-injection approach.
The analysis shows that LLMs typically:
-
Begin by
solving the problem correctly first
to establish an anchor. -
Articulate and simulate multiple potential misconceptions
through erroneous reasoning. -
Perform plausibility assessments and curate a final set of distractors.
In contrast, similarity-based strategies are only rarely observed,
accounting for approximately 3% of samples.
Diagnosing failure modes
The authors use their taxonomy to determine where the distractor generation pipeline fails. Their analysis reveals that errors do not primarily arise from failures in simulating errors or structuring the process.
Instead, failures are attributed to two main issues:
-
Failures in recovering the correct solution,
which propagate downstream and degrade subsequent generation. -
Errors in
selecting among response candidates,
specifically regarding inaccurate plausibility assessment or curation.
The research highlights that providing the correct solution in the prompt improves alignment with human-authored distractors by 8%.
This finding emphasizes the critical role of anchoring to the correct solution
when attempting to generate plausible, incorrect student reasoning. Overall, while LLMs can systematically model misconception-based generation, their performance is heavily contingent on successful initial problem-solving.
Improvements for AI systems
1. Anchored Dual-Stage Reasoning Pipeline
-
Improvement: Implement a mandatory two-stage execution architecture where the model is prohibited from generating distractors until a verified correct solution is established and passed as a hard constraint to the second stage.
-
Capability: The system will eliminate
error propagation
failures by ensuring that all subsequent misconception simulations are branched from a mathematically soundanchor
solution, preventing the generation of distractors based on incorrect fundamental math.
2. Misconception-Based Over-generation and Discriminative Ranking (OBDR) Module
-
Improvement: Replace single-shot distractor generation with an iterative
over-generate and rank
architecture. The system will generate a large pool of N candidate distractors derived from diverse error descriptions (ERR DESC), followed by a secondary agentic pass to perform plausibility assessment (PLAUS) and curation (CURATE). -
Capability: The system will move beyond simple similarity-based heuristics (e.g., sign flipping) to produce high-quality, diagnostically informative distractors that specifically target identifiable procedural gaps in student reasoning.
3. Taxonomy-Guided Chain-of-Thought (TG-CoT) Fine-tuning
-
Improvement: Fine-tune LLMs using a structured training objective based on the paper's taxonomy, specifically enforcing the sequence: ERR DESC to ERR SIM to INST.
-
Capability: The system will force the model to explicitly verbalize the abstract logic of a misconception (the
why
) before performing the erroneous calculation (thehow
), ensuring that generated distractors are logically consistent with actual human error patterns rather than random mathematical hallucinations.
4. Multi-Agent Discriminative Verification Gate
-
Improvement: Integrate a specialized
Plausibility Evaluator
agent that uses a discriminative scoring function to rank candidate distractors based on their predicted ability to attract low-performing students while maintaining zero overlap with the correct answer. -
Capability: The system will automatically filter out
non-functional
distractors (those that are too obvious or mathematically impossible) and prioritize those with high negative discrimination, significantly increasing the utility of generated MCQs for diagnostic assessment.
Abstract
Modeling student misconceptions in a realistic manner is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating distractor answers for multiple-choice questions (MCQs), a task that requires producing answers that are incorrect, yet plausible. We introduce a taxonomy over reasoning strategies for distractor generation that is grounded in learning-science literature and empirical observation, which we apply to LLM-generated reasoning traces across math and science MCQs. On the math dataset, we find that models follow a misconception-based process with potentially high diagnostic value: they recover the correct solution, articulate student errors, simulate them, and select plausible candidates. On the science dataset, on the other hand, they tend to follow a less robust approach based on semantic similarity to the correct answer. We find the most frequent failure modes to be that the model is unable to generate a correct solution or that it discards plausible distractor candidates when performing selection. Providing the correct solution in the prompt yields a relative improvement of 6.4% in alignment with human-authored distractors, highlighting the critical role of anchoring distractor generation to the correct solution. Together, our findings offer an interpretable view of how LLMs model incorrect student reasoning.
Sources
- Math Multiple Choice Question Generation via Human-Large Language Model Collaboration
- Novice Learner and Expert Tutor: Evaluating Math Reasoning Abilities of Large Language Models with Misconceptions
- DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
- Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering