INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators

arXiv:2608.10492 · cs.AI, cs.CY · Submitted 2026-08-14 · Read on arXiv

Rose Niousha, Minwoo Kang, Narges Norouzi

University of California, Berkeley

cs.AI, cs.CY

Submitted: 2026-08-14

Updated: 2026-08-17

Comments: Accepted at the Conference on Language Modeling (COLM) 2026

Code: https://github.com/rosensh/inside

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper presents INSIDE (INTERNAL STUDENT DIALOGUE), a student modeling framework that fine-tunes Large Language Models (LLMs) to simulate both the observable actions and the latent internal

Terminology

Summary

The paper presents INSIDE (INTERNAL STUDENT DIALOGUE), a student modeling framework that fine-tunes Large Language Models (LLMs) to simulate both the observable actions and the latent internal reasoning of students in educational settings. The work addresses a critical gap in LLM-based student simulation: while existing approaches can reproduce observable student behaviors (like code submissions), they fail to capture the underlying reasoning processes that drive those behaviors. As the paper states, Two students may submit identical submissions for entirely different reasons.

The paper identifies that LLM-based simulations are often limited to replicating surface-level patterns of user behavior, rather than modeling the latent processes underpinning observable outcomes. In educational contexts, understanding the reasoning behind observed actions is crucial for applications such as diagnosing misconceptions, generating targeted feedback, and evaluating tutoring systems. The authors note that existing work on LLM reasoning has largely focused on improving correctness, encouraging models to produce logically consistent or factually accurate outputs, but human actions, however, are not always rational or correct: people frequently make errors, hold misconceptions, or apply incomplete strategies.

The paper makes three main contributions:

  1. Reconstruction of pedagogically grounded reasoning traces: The authors reconstruct latent reasoning traces from student interaction data via retrospective inference with a teacher model, conditioning on prior context and observed code edits. These traces are grounded in educational theory and enable models to generate interpretable internal dialogue despite the absence of ground-truth reasoning.

  2. A student modeling framework with internal dialogue: INSIDE jointly models student code generation and the underlying reasoning process by conditioning actions on inferred internal dialogue.

  3. A two-dimensional evaluation of simulation fidelity: The framework is evaluated on "(1) action fidelity, the similarity between generated and real student code, and (2) reasoning quality, defined as the alignment between generated reasoning and ground-truth code edits without requiring observed reasoning traces."

The study uses data from an introductory programming course at the University of California, Berkeley, with ∼900 students per semester. The data spans two semesters (Spring 2024 and Spring 2025) focusing on the first five homework assignments about Python programming. The training set uses Spring 2025 data (445 students, 2,022 submission streams, 6,911 total submissions), and the test set uses Spring 2024 data (479 students, 1,546 submission streams, 6,316 total submissions).

Two test subsets are defined:

  • test OP (Old Problems): 5,262 code submissions from new students on problems that appear in the training set

  • test NP (New Problems): 1,054 code submissions from students solving problems not present in the training data

The framework generates internal dialogue through a two-stage process using a teacher model (GPT-5). First, the teacher LLM infers the student's internal state in third person, producing structured summaries of each cognitive, affective, and action states, inspired by three domains in Bloom's Taxonomy. Then, given these inferred states, the model generates a first-person internal dialogue (a think trace) reflecting the student's reasoning process prior to the observed submission.

The interaction context at time step t is defined as: xt = pu, c<t,si,pu, f<t,si,pu, where pu is the problem, c<t are prior code submissions, and f<t are prior feedback instances. The model generates internal dialogue zt and next code submission ct: zt,si,pu ∼ M(· xt) and ct,si,pu ∼ M(· zt,si,pu, xt).

Two experiments are conducted:

  • Experiment 1 (without CoT): Generate code submission conditioned on prior submissions and feedback, without intermediate reasoning generation

  • Experiment 2 (with CoT): Generate both internal dialogue and code submission as a single sequence

Models evaluated include:

  • Fine-tuning: Qwen2.5-7B, Qwen2.5-Coder-7B, Qwen3-8B-Base, and LLaMA-3-8B using LoRA (r=16, α=32), for two epochs with learning rate 10−4

  • Prompting: Instruction-tuned variants including Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Qwen3-8B, LLaMA-3-8B-Instruct, and GPT-5, plus reasoning models Qwen3-14B and Qwen3-32B

Naming conventions: model name-SFT (Experiment 1 fine-tuning), model name-INSIDE (Experiment 2 fine-tuning), model name-CoT (standard CoT prompting), model name-BloomCoT (Bloom-inspired structured CoT prompting).

Action fidelity is evaluated by comparing model-generated code against real student submissions using:

  • Code functionality: Pass rate on the autograder test suite

  • Code complexity and style: Lines of code (LOC), AST depth and width, PEP 8 violations

Wasserstein distance (Earth Mover's Distance) is reported for each metric, where lower values indicate closer alignment between distributions.

The alignment metric evaluates whether the generated internal dialogue reflects the real student's code changes. An LLM judge (GPT-5-mini) decomposes the synthetic internal dialogue into a set of atomic claims Vt = v(1),..., v(n) representing intended actions. Each claim is evaluated against the ground-truth diff, and alignment is computed as the fraction of supported claims.

The judge validation shows an average alignment score of 95.2% for transitions ct−1,si,pu → ct,si,pu given zt,si,pu, and manual annotations achieved 88.0% agreement with the LLM judge labels (κ = 0.754).

On test OP, INSIDE consistently achieves the lowest Wasserstein distances across all metrics, indicating the closest match to real student code distributions. Compared to SFT baselines, incorporating internal dialogue improves alignment not only in functionality but also in stylistic and structural properties of code.

On test NP, results are more mixed: models fine-tuned with regular SFT and INSIDE perform comparably across most metrics. The paper explains this asymmetry: test NP contains a higher proportion of successful student submissions than test OP... When student failures are more common, as in test OP, this remaining mismatch is larger, so INSIDE has more room to improve.

Pass rate trajectory analysis shows that "real students start with low pass rates and exhibit a sharp increase near the final steps, reflecting incremental progress. SFT models closely track this trajectory. In contrast, prompting-based models maintain relatively high pass rates (≈ 80%) from the beginning, with little variation across steps."

On test OP 1, the top models by MAE are Qwen3-8B-INSIDE (0.094), Qwen2.5-Coder-7B-INSIDE (0.098), and Qwen2.5-7B-INSIDE (0.113).

INSIDE achieves the highest alignment across both settings:

  • On test OP: Qwen2.5-7B-INSIDE reaches 51.8%, outperforming the best prompting baseline, Qwen2.5-7B-Instruct-BloomCoT

  • On test NP: Qwen3-8B-INSIDE achieves 57.9% compared to 56.0% for the strongest BloomCoT model

Interestingly, "larger and more capable models (e.g., GPT-5 and Qwen3-32B) tend to achieve lower alignment scores, suggesting that stronger reasoning ability does not necessarily translate to reasoning that matches student-like code edits."

The paper emphasizes that alignment should be interpreted jointly with action fidelity. High alignment alone does not indicate realistic student modeling if the generated code does not follow plausible solution trajectories. While some prompting models show comparable alignment, these models exhibit poor action fidelity and unrealistic solution trajectories, indicating a disconnect between explanation and behavior. In contrast, INSIDE achieves both high alignment and strong action fidelity.

Prompting-based models achieve high self-consistency scores (86.9%–99.0%), with GPT-5 attaining the highest. However, GPT-5 performs among the worst in terms of reproducing realistic student code trajectories. INSIDE achieves competitive self-consistency (83.0%–87.3%), indicating that its generated reasoning remains coherent with its generated code while better matching real student code edits.

The paper acknowledges several limitations:

  1. Reconstructed reasoning: Internal dialogue is reconstructed rather than observed. The reasoning traces are generated by a teacher LLM through retrospective inference, and therefore represent an approximation of the student's latent cognition.

  2. Expert-like reasoning bias: Because LLMs are typically trained to produce expert-like reasoning, they may struggle to faithfully reconstruct novice reasoning patterns, even in a reconstruction setting.

  3. Distributional limitation: The splits between test OP and test NP differ not only in whether problems are seen or unseen, but also in their underlying student pass-rate distributions.

  4. Alignment ceiling: While INSIDE achieves the highest alignment among models (reaching ∼58%), this shows that a majority of generated claims explain student code edits, with remaining gaps indicating room for further improvement.

The paper concludes that modeling reasoning beyond observable actions is key to building more realistic student simulators and enables new opportunities for evaluating and improving tutoring systems. INSIDE points to LLM-based simulations grounded in cognitively plausible internal reasoning, presenting opportunities for developing tutoring systems equipped with misconception-aware interventions. The framework enables new applications including evaluating whether feedback resolves misconceptions and supports counterfactual analysis of alternative interventions, as well as learner-facing tools that can support reflection and metacognition.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems:

Improvement: I can generate both observable student actions (code submissions) AND latent internal reasoning (think-aloud traces) that are grounded in educational theory (Bloom's Taxonomy cognitive, affective, and action states).

What the improved system can do:

  • Simulate a student who submits incorrect code because of a specific misconception (e.g., confusing = with ==), not just because it's wrong

  • Distinguish between two students who submit identical code but for different reasons (one forgot a syntax rule, another misunderstood loop semantics)

  • Generate realistic novice reasoning that includes errors, partial understanding, and incomplete strategies—not just expert-like correct reasoning

Abstract

Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.

Sources

Related papers