Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank
Joshua Mitton, Prarthana Bhattacharyya, Digory Smith, Thomas Christie, Ralph Abboud, Simon Woodhead
Eedi · Renaissance Philanthropy · Learning Engineering Virtual Institute
cs.CL, cs.LG
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 21 pages, 8 figures, 8 tables. Joshua Mitton and Prarthana Bhattacharyya contributed equally to this paper
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
Terminology
Summary
Summary
This paper presents a novel approach for detecting student misconceptions from student-tutor dialogues using large language models (LLMs). The authors introduce a three-stage pipeline—generation, retrieval, and reranking—to identify the most likely misconception a student holds, given a multiple-choice math question, the student's chosen answer, and a dialogue between the student and a tutor.
Problem and Motivation: The paper addresses the challenge of timely and accurate identification of student misconceptions, which is critical for improving learning outcomes and preventing the compounding of errors. The authors note that misconceptions are ambiguous to define, often overlap, and are difficult to detect from limited student activity. Traditional knowledge tracing models trained on answer correctness alone often peak at around 65-70% accuracy. The paper argues that student-tutor dialogues provide richer context that can improve misconception diagnosis, but introduce the complexity of mapping dialogue to generalizable, yet specific, tutor-defined misconception labels.
Proposed Model: The proposed model consists of three stages, as illustrated in Figure 1:
-
Stage 1 (Generation): A fine-tuned LLM generates a plausible misconception hypothesis from the input data (question text, answer text, and student-tutor dialogue). The model uses parameter-efficient LoRA adapters for fine-tuning.
-
Stage 2 (Retrieval): An off-the-shelf MiniLM-L6-v2 embedding model embeds the generated misconception and each ground-truth misconception label. The top-k misconceptions are retrieved by ranking based on cosine similarity between the generated and ground-truth misconception embeddings.
-
Stage 3 (Reranking): A fine-tuned LLM reranks the top-k misconception labels to better capture the semantics of the misconceptions. The highest-ranked misconception label is the final prediction.
Dataset: The authors collected a new tutor-student dialogue dataset from an online educational learning platform. Human tutors labelled the likelihood that a student has a specific misconception on a scale of 0, 25, 50, 75, 100. Only data points where the tutor labelled the likelihood as 75 or 100 were retained, resulting in 922 student-tutor conversations. The dataset contains 546 unique ground-truth misconception labels, with the majority of misconceptions occurring only once. The dataset was split by unique misconception labels (70% training, 10% validation, 20% testing) to ensure that misconception targets are unique between datasets and to assess model generalization to unseen misconceptions.
Experimental Protocol: The models were evaluated using several metrics: MAP@k (Mean Average Precision at k), NDCG (Normalized Discounted Cumulative Gain), recall@k, cosine similarity between generated and true misconception labels, and mean and median retrieval rank.
Baselines: Three baselines were compared:
-
Direct Embedding Matching: Skips the LLM generation step, directly embedding student dialogue text using MiniLM-L6-v2 and matching against misconception labels via cosine similarity.
-
Zero-shot LLM Classification: Presents Claude Sonnet 4.5 with the student dialogue and the complete list of 546 unique misconception labels, prompting it to directly rank the top-10 most likely misconceptions in a single inference step.
-
TF-IDF Keyword Matching: Uses TF-IDF vectors and cosine similarity to rank misconceptions based on keyword overlap.
Base LLM Models: The authors evaluated both closed-source (Claude Sonnet 4.5) and open-source models (Llama 3.2 3B Instruct and Qwen 2.5 7B Instruct). They also fine-tuned the open-source models using Low-Rank Adaptation (LoRA), keeping only about 0.5% of parameters trainable.
Key Results:
-
Baselines: Keyword matching outperforms zero-shot LLM classification and direct embedding on metrics prioritizing top-1 prediction, while direct embedding is stronger for recall@k. The zero-shot LLM classification baseline scores poorly due to struggling with the large context of 546 unique misconceptions.
-
Base Model Comparison: Llama 3.2 3B is significantly outperformed by Claude Sonnet 4.5 and Qwen 2.5 7B, particularly in reranking. Qwen 2.5 7B outperforms Claude on MAP@k, recall@3, and median rank, while Claude outperforms Qwen on NDCG, recall@10, and mean rank. Overall, Claude and Qwen perform comparably despite Qwen being significantly smaller.
-
Ablation of Stages: The generative stage improves overall model performance for fine-tuned Llama 3.2 3B LoRA, zero-shot Qwen 2.5 7B, and fine-tuned Qwen 2.5 7B LoRA. Reranking improves final predictions for Claude Sonnet 4.5 and Qwen 2.5 7B across all metrics, but hurts zero-shot Llama 3.2 3B performance. Fine-tuning improves Llama 3.2 3B LoRA across all metrics and generally improves Qwen 2.5 7B, notably improving NDCG so that it outperforms Claude.
Discussion:
-
Base Model Selection: Models with stronger zero-shot reasoning capabilities (Claude, Qwen) produce more accurate misconception descriptions than Llama 3.2 3B. Qwen 2.5 7B performs comparably to Claude despite being much smaller, suggesting domain-relevant pre-training matters more than raw parameter count.
-
Fine-tuning: Fine-tuning with LoRA improves generated misconception style, producing outputs that are more concise and stylistically closer to expert-authored labels. Table 1 shows that zero-shot outputs are verbose, while fine-tuned outputs match tutor format. Fine-tuning improves cosine similarity between generated outputs and expert labels.
-
Semantic Similarity vs. Mathematical Coherence: Embedding-based approaches can match misconceptions stylistically but may differ mathematically. Table 2 presents examples where predicted misconceptions share vocabulary/domain but differ in underlying mathematical reasoning, highlighting the need for math-centric language models and better retrieval metrics.
-
Reranking: Reranking with LLMs improves rank-1 accuracy by understanding semantic relationships beyond surface-level word matching. Table 3 shows examples where reranking promoted the correct misconception from rank 2 or lower to rank 1. Fine-tuning dedicated reranking models improves performance, and smaller specialized models can outperform larger closed-source models.
Conclusion: The authors present a generation-retrieval-reranking pipeline that outperforms baselines in misconception diagnosis from dialogue. They develop a parameter-efficient LoRA fine-tuning approach that improves generated misconception style and can outperform closed-source models an order of magnitude larger. Extensive ablations and benchmarking demonstrate the importance of each proposed component.
Additional Details from Appendix:
-
The appendix provides additional metrics (mean and median rank) and results.
-
Prompt engineering experiments showed that a
With Examples
strategy (including positive and negative examples, with explicit length constraints) achieved a 12% improvement in semantic similarity and 84% reduction in word count compared to a naive verbose prompt. -
Table 7 shows that reranking successfully promoted the correct misconception to rank 1 in 73 cases for Claude Sonnet 4.5, with improvements from various baseline ranks (2-10).
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:
Improvement: Implement a generate-retrieve-rerank architecture instead of single-pass classification.
What the improved system can do:
-
Stage 1: Generate a concise, hypothesis-style misconception description from student-tutor dialogue (not a direct label prediction)
-
Stage 2: Retrieve top-k candidate misconceptions from a predefined taxonomy using embedding similarity (MiniLM-L6-v2)
-
Stage 3: Rerank candidates with a fine-tuned LLM to capture semantic nuances beyond lexical overlap
Result: Achieves 0.44 MAP@10 and 0.43 NDCG with Qwen 2.5 7B LoRA fine-tuned, outperforming zero-shot Claude Sonnet 4.5 (0.42 MAP@10, 0.42 NDCG) despite being 10x smaller.
Abstract
Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors. However, this task is highly dependent on the effort and intuition of the teacher. In this work, we present a novel approach for detecting misconceptions from student-tutor dialogues using large language models (LLMs). First, we use a fine-tuned LLM to generate plausible misconceptions, and then retrieve the most promising candidates among these using embedding similarity with the input dialogue. These candidates are then assessed and re-ranked by another fine-tuned LLM to improve misconception relevance. Empirically, we evaluate our system on real dialogues from an educational tutoring platform. We consider multiple base LLM models including LLaMA, Qwen and Claude on zero-shot and fine-tuned settings. We find that our approach improves predictive performance over baseline models and that fine-tuning improves both generated misconception quality and can outperform larger closed-source models. Finally, we conduct ablation studies to both validate the importance of our generation and reranking steps on misconception generation quality.
Sources
- DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
- MiRAGE: Misconception Detection with Retrieval-Guided Multi-Stage Reasoning and Ensemble Fusion
- The Llama 3 Herd of Models
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Toward In-Context Teaching: Adapting Examples to Students' Misconceptions
- Learning to Make MISTAKEs: Modeling Incorrect Student Thinking And Key Errors
- Multilingual E5 Text Embeddings: A Technical Report
- Instructions and Guide for Diagnostic Questions: The NeurIPS 2020 Education Challenge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering