EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations".
Jane: The paper was written by Baptiste Moreau-Pernet from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the channel, folks! We’ve got a fascinating paper on our hands today, and I’m thrilled to be digging into it with Jane. It’s called “EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations.” Jane, I’ve got to say, that title alone is a mouthful, but it’s pointing at something really important.
Jane: Oh, absolutely, Tom! And I love that the title spells out exactly what it’s about. We’re talking about simulating learners — you know, AI students that can practice conversations with a tutor — and then figuring out whether those simulations actually behave like real students. The author, Baptiste Moreau-Pernet, is trying to solve a problem that’s been bugging a lot of people in education tech.
Tom: Right, and I think the key word there is “evaluating.” We’ve had simulated learners for years, but how do we know if they’re any good? That’s the gap this paper is trying to fill. Jane, can you break down what a learner simulation even is for our listeners who might be new to this?
Jane: Sure! Imagine you’re building an AI tutor, like the ones that help kids with math homework. You want to test that tutor before you unleash it on real students, right? So you create a simulated student — an AI that pretends to be a learner, makes mistakes, asks questions, and learns over time. That way, you can see how your tutor handles different situations without needing hundreds of real kids to test it on.
Tom: And that’s where the problem comes in. These simulated learners are powered by large language models now, and they can be super realistic in some ways, but they’re also trained to be helpful assistants, not necessarily to mimic how a struggling student actually talks. So the question becomes: how do you know if your simulated student is actually acting like a real one?
Jane: Exactly! And that’s what EvalConvoLearn is all about. It’s an open-source framework, which means anyone can use it and build on it. The author is giving the community a way to measure how close a simulated learner gets to real student behavior. And I think that’s a big deal because it’s the first open framework of its kind.
Tom: I love that it’s open-source, too. There’s been similar work before, but it was closed-source, so researchers couldn’t easily reproduce the results or adapt it to their own datasets. This paper is really about democratizing that evaluation process.
Jane: And the implications are huge, Tom. If we can reliably evaluate these simulations, we can use them to test educational theories, improve AI tutors, and even create teachable agents where real students learn by teaching a simulated peer. That’s a whole new way to think about education.
Tom: I’m already excited about where this is going. But before we get ahead of ourselves, let’s talk about what the framework actually does in the next segment. Jane, you’re going to love the details.
Jane: I can’t wait, Tom. Let’s dig into the summary and see how EvalConvoLearn measures all of this.
Summary: Tom: So, Jane, we’ve set the stage. Now let’s get into the meat of “EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations.” What does this framework actually measure?
Jane: Great question, Tom. The framework evaluates simulations along two main axes. First, there’s learning behavior — that’s whether the simulated student actually masters skills over time the way a real student would. Second, there’s conversational quality — things like how often the student asks questions, how long their responses are, and what kinds of errors they make.
Tom: And that’s not just a vibe check, right? They’re grounding these metrics in real data. They use a dataset of actual tutoring dialogues from a platform called Eedi, which is publicly available. So they’re comparing the simulated learner’s behavior to what real students actually did in those conversations.
Jane: Exactly! And here’s a clever part: they don’t just run one generic conversation. They create different “scenarios” based on the skill being practiced and the learner’s prior mastery. So you might have a scenario where the student already knows the prerequisites but hasn’t mastered the target skill yet. That way, you’re testing the simulation in specific, realistic situations.
Tom: And they cap conversations at seven turns, which I thought was interesting. They’re using a heuristic from learning science that says most students need about seven attempts to master a skill. So that keeps the simulation realistic without dragging on forever.
Jane: Right, and then they score the simulation on how close it gets to real student distributions. For learning behavior, they use something called L1 distance — basically, how different are the simulated mastery outcomes from the real ones? For conversational quality, they use Jensen-Shannon divergence and Wasserstein distance, which are fancier ways of measuring how similar two distributions are.
Tom: And here’s the kicker — they tested two different simulated learners. One keeps summaries of past conversations, and the other tracks skills as binary mastered or not mastered. Both are powered by GPT-four point one-mini, and they found that both learners solve about ninety percent of the items, which is way more than real students. Real students only solved about thirty-four percent of the items in that dataset.
Jane: That’s a huge gap, Tom! It tells us that these simulations are too good at solving problems. They’re not making the kinds of mistakes real students make. And that’s exactly the kind of insight EvalConvoLearn is designed to surface — it gives you a signal to iterate on your simulation design.
Tom: And I love that they also anchored the tutor responses in real tutor utterances from the dataset. So the tutor isn’t just some generic AI — it’s grounded in how actual human tutors talk. That makes the whole evaluation more realistic.
Jane: It really does. And they even did a small validation of their labeling — they had humans check the AI’s labels for talk moves and error types, and the agreement was pretty high. So the metrics aren’t just coming out of thin air.
Tom: Alright, so we’ve got the framework and the results. But what does this mean for the future? What improvements are they suggesting? Let’s get into that next.
Jane: Good segue, Tom. Let’s talk about where this is heading.
Improvements and Future Work: Tom: So, Jane, we’ve seen how EvalConvoLearn works and what it found. But the paper doesn’t stop there — it’s very much a work in progress, and the author is pretty clear about what needs to improve. What stood out to you?
Jane: Well, the biggest thing for me is that they want to expand the suite of scenarios. Right now, they’re working with a pretty small dataset — only sixty-six conversations from the Eedi platform. That’s enough to demonstrate the framework, but it’s not enough to make broad claims about learner simulations in general. They need more data, more skills, more diverse tutoring contexts.
Tom: And they’re also looking at the tutor side of things. In the appendix, they did an ablation study where they compared the tutor with and without few-shot prompting. Adding real tutor examples helped the learning behavior scores, but it actually made the conversational scores worse for one of the learners. That’s a fascinating trade-off.
Jane: Yeah, that really caught my eye too. When the tutor used three-shot prompting — meaning it saw real tutor responses as examples — the simulated learners were less likely to solve problems in seven turns. That makes sense because real tutors don’t just give away the answer; they ask guiding questions. So the simulations became more realistic in terms of learning behavior, but the error types shifted in ways that didn’t match real students as well.
Tom: So it’s a balancing act. You can’t just tune the learner in isolation — you have to tune the tutor and the learner together. The author even calls it a “dual learner-tutor optimization task.” That’s a really important insight for anyone building these systems.
Jane: And they’re also thinking about validation. They want to get human ratings on the metrics to make sure they’re actually capturing what matters. Right now, the metrics are automated, which is great for scalability, but you need that human check to make sure you’re measuring the right things.
Tom: I also noticed they mentioned fine-tuned or RL-trained models as future test cases. So instead of just using off-the-shelf LLMs, you could train a learner simulation specifically to match real student behavior, and then use EvalConvoLearn to see how well it does. That could be a real breakthrough.
Jane: Absolutely. And the open-source nature of the framework means researchers can adapt it to their own datasets and add new metrics. It’s not a one-size-fits-all tool — it’s a foundation that the community can build on.
Tom: So, what’s the big picture here? If this framework gets adopted, what does it mean for education technology?
Jane: I think it means we can finally start benchmarking learner simulations the way we benchmark other AI systems. That’s how we make progress — you need a standard way to measure performance. And once we have that, we can build simulations that are actually useful for testing tutors, for research, and maybe even for students who learn by teaching.
Tom: I’m with you on that. Let’s wrap this up in the conclusion and talk about why this paper matters for the world.
Conclusion: Tom: Alright, Jane, we’ve covered a lot of ground on “EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations.” Let’s bring it home. What’s the one thing you want our listeners to remember?
Jane: I think it’s that we finally have an open, standardized way to ask whether a simulated learner is actually acting like a real student. That’s not a trivial question — it’s the difference between building AI tutors that work in theory and building ones that work in practice. EvalConvoLearn gives us a yardstick.
Tom: And the results show there’s a long way to go. These simulations are too good at solving problems — they don’t make enough mistakes. But that’s not a failure; that’s a signal. It tells us exactly where to focus our efforts.
Jane: Right, and the framework is designed to be extended. More scenarios, more metrics, more datasets. The author is inviting the community to come in and build on this work. That’s how you make real progress.
Tom: I also love that they’re thinking about the tutor and the learner as a pair. You can’t just tune one side — you have to think about how they interact. That’s a lesson that goes beyond education, honestly.
Jane: It really does. And the fact that it’s open-source means anyone can pick it up and start using it tomorrow. That’s the kind of work that moves the field forward.
Tom: Well said, Jane. We’ve had a great time digging into this paper, and I hope our listeners feel the same. We’ll be back with another exciting paper soon, so stay tuned!
Jane: Thanks for joining us, everyone. Until next time, keep learning — and maybe teach an AI something while you’re at it!
Baptiste Moreau-Pernet
cs.HC, cs.CL
Submitted: 2026-06-18
Comments: Poster at the Impactful and Responsible AI Systems for Education workshop, as part of the Festival of Learning 2026
Code: https://github.com/RenaissancePhilanthropy/EvalConvoLearn
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 64/100
The gist: "We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves,
Terminology
Summary
Summary
The paper introduces EvalConvoLearn, an open-source framework for evaluating conversational learner simulations in tutoring contexts. The authors state: "We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). The framework
measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances."
The motivation is that "LLM-based tutors interact with students through naturalistic conversations (e.g., Khan Academy's Khanmigo, Google LearnLM). Evaluating such tutors at scale requires correspondingly capable conversational learner simulations. However,
ensuring that such simulations are valid and useful for the various goals above is an open problem: simulating realistic student learning and conversational characteristics is complex because LLMs are trained as helpful assistants. Prior work by Scarlatos et al. (2025)
introduce a consolidated set of metrics by matching a simulated learner response with the real student message given the conversation history. However, they only evaluate simulations on one turn and not at the dialogue level, and their work is closed-source, making reproduction and extensions to other datasets difficult."
The framework works as follows: "Given a tutoring conversation dataset with practice items and tagged skills (or items automatically tagged), EvalConvoLearn extracts learning scenarios defined as the product of a target skill and the learner's prior mastery of it. For each scenario, the framework runs simulated problem-solving conversations between the custom learner and EvalConvoLearn's tutor, then compares simulated and real metric distributions. The tutor is grounded at the individual tutor level with few-shot dialogue examples from other conversations in the dataset to simulate more realistic responses."
The learner simulation interface requires four functions: "The learner initializes its knowledge state according to a skill profile (mastered/unmastered skills or prerequisites). Then, it generates a response given a conversation history and its knowledge state, and is able to update its knowledge state from a conversation. Finally, it should be able to return whether it currently masters a given skill. If a custom learner cannot look up skills,
a default initialization loop runs problem-solving conversations along the skill's prerequisites tree, while assessing progress at each step via skill-aligned items using the learner's response generation functionality."
For scenario-based evaluation, we treat each learner turn as a skill practice attempt. We cap conversations at 7 turns to allow time for learning while preventing unrealistic lengths, following a heuristic of seven attempts to mastery.
The learning-behavior score LB is the L1 distance between simulated and real outcome distributions for each scenario and skill, macro-averaged across the skill pool.
Conversational realism is evaluated with four metrics: The presence of three talk moves and five error types (LLM-labeled), number of questions in interrogative learner turns, and turn length.
The authors report We manually evaluated LLM labels on 50 conversations and got a per-label Cohen's κ between.73 and 1.0 for talk moves (exact match.96) and from.79 to 1.0 for error types (exact match.88).
Each metric is scored by computing the distance between its distribution in the simulated and real data within each scenario, using Jensen–Shannon divergence (JSD) or Wasserstein-1 distance.
The conversational score CONVs averages the four distance metrics with equal weights.
The final EvalConvoLearn score (ECL) averages CONVs and LBs across scenarios s, accounting for scenario distributions from their occurrence in real dialogues.
The framework is evaluated on the largest publicly available dataset of student–TA tutoring dialogues, from the Eedi learning platform.
The authors assume students proactively seek help on Eedi, we assume they have not yet mastered the target item skill but have mastered its prerequisites given Eedi's progressive curriculum.
They retain conversations with ≥50% student utterances, conducted by tutors with ≥5 conversations (to enable few-shot prompting), yielding 66 conversations.
Each practice item is tagged with a single skill from a subset of the B.E.S.T. curriculum using GPT-4.1-mini.
Two LLM-based learners are tested: "The Conversation Summaries learner stores summaries of past conversations and uses them as context when generating a new response. The Binary Skills learner uses LLM-as-a-judge to tag items with skills and stores binary mastery outcomes after conversations. Results are
aggregated across 3 skills with 8 conversations per skill, using
Claude's Sonnet-4.6 for tutor responses and metric evaluations."
Results show: The learners have similar LB scores as they both solve around 90% of items, which is 56 points more than real students, but differ slightly in conversational metrics while keeping high within-learner consistency.
Specifically, the Binary Skills learner achieved CONV of 0.433 ± 0.011, LB of 0.556 ± 0.014, and ECL of 0.494 ± 0.009; the Conversation Summaries learner achieved CONV of 0.450 ± 0.004, LB of 0.556 ± 0.014, and ECL of 0.503 ± 0.008. The authors note These results provide a signal to iterate on each learner's design by targeting their conversational or learning abilities, and suggest a large room for improvement for the models to produce realistic learner behavior.
An ablation study on tutor prompting found: "adding few-shot prompting improves final ECL scores for both learners, mainly through consistently better learner behavior. However, conversational scores see no change or get worse (in the case of the Binary Skills learner) when adding few-shot prompting. The authors hypothesize
Adding human tutor examples seems to reduce the propensity of learners to solve problems in 7 conversation turns, which may be explained by shorter, more conversational, or less information-rich tutor responses. They note
Simple analysis (N=450) corroborates this hypothesis by showing longer average tutor responses in the 0-shot setting (51 words) compared to the 3-shot setting (46 words)."
Future work includes: "we are looking to expand the suite of scenarios and further validate the metrics with human ratings. When released, users can test the framework with new datasets, advanced learner and tutor simulations (e.g. fine-tuned or RL-trained models), and iterate on the learner quality signal to develop learner simulation benchmarks for various educational contexts. The paper also mentions
moving towards a dual learner-tutor optimization task."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:
-
Improvement: Replace generic, verbose tutor responses with few-shot prompting using real tutor utterances from the dataset (3-shot examples from other conversations).
-
Capability: The AI tutor now produces shorter, more conversational, and contextually appropriate responses that mimic human tutor behavior, reducing the tendency to over-explain or provide excessive information.
-
Improvement: Implement a two-axis evaluation framework (learning behavior + conversational quality) that compares simulated learner distributions against real student data using JSD and Wasserstein-1 distances.
-
Capability: The system can now quantitatively assess whether a simulated learner is realistically acquiring skills and conversing, rather than relying on surface-level plausibility.
-
Improvement: Add a binary skill mastery state that updates after each conversation, with an LLM-as-a-judge to tag items with skills.
-
Capability: The AI can now track which skills a simulated learner has mastered and generate responses conditioned on that knowledge state, enabling more realistic learning trajectories.
-
Improvement: Store and retrieve summaries of past conversations as context for generating new learner responses.
-
Capability: The system can maintain coherent, long-term learner behavior across multiple interactions, avoiding contradictions and improving conversational consistency.
-
Improvement: Use LLM-based labeling (with validated Cohen’s κ between.73–1.0) to classify learner utterances into three talk moves and five error types.
-
Capability: The AI can now automatically detect and categorize student discourse patterns and misconceptions, enabling fine-grained analysis of conversational realism.
-
Improvement: Add quantitative metrics for turn length (in words) and question frequency in interrogative turns.
-
Capability: The system can now measure whether simulated learners ask questions at realistic rates and produce appropriately sized responses, matching human student behavior.
-
Improvement: Implement an ablation study that compares 0-shot vs. 3-shot tutor prompting to isolate the effect of tutor behavior on learner scores.
-
Capability: The system can now identify whether differences in learner performance are due to learner design or tutor behavior, enabling more targeted improvements.
-
Improvement: Weight final ECL scores by the frequency of each scenario in real dialogues.
-
Capability: The evaluation now prioritizes common learning scenarios, ensuring that improvements in rare edge cases don’t overshadow performance in typical tutoring interactions.
-
Benchmark learner simulations across different architectures (e.g., binary skill vs. conversation summary) with reproducible, open-source metrics.
-
Identify specific weaknesses in a simulated learner—e.g., whether it fails on conversational realism (high CONV score) or learning behavior (high LB score)—and guide targeted retraining.
-
Generate more human-like tutor responses that are shorter, more conversational, and grounded in real tutoring data, reducing the “verbose assistant” bias of LLMs.
-
Track skill mastery over time and simulate realistic learning curves, including partial mastery and prerequisite-based progression.
-
Automatically classify student talk moves and error types with high reliability, enabling large-scale analysis of tutoring dialogues without manual annotation.
-
Compare tutor models (e.g., Sonnet 4.6 vs. others) to determine which produces more realistic learner behavior, supporting dual learner-tutor optimization.
-
Adapt to new datasets by plugging in custom learner interfaces and scenario definitions, making it a general-purpose evaluation tool for educational AI.
These improvements directly address the paper’s identified gaps: open-source evaluation, dialogue-level (not just turn-level) assessment, and grounding in real student data. The resulting system is more rigorous, reproducible, and actionable for improving both simulated learners and AI tutors.
Abstract
Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.
Sources
- Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
- TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
- Towards Valid Student Simulation with Large Language Models
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support