Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients

arXiv:2608.07499 · cs.HC, cs.AI, cs.CL · Submitted 2026-06-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients".

Jane: The paper was written by Jiading Zhu, Xinyu Cindy Wang, Thomas Nguyen, Yan Qing Lee, Osnat C. Melamed et al. from University of Toronto and Centre for Addiction and Mental Health.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's got a real mouthful of a title: "Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients."

Jane: And Tom, I have to say, that title is dense, but what's underneath it is genuinely fascinating. It's about using AI to train and evaluate counselors who help people quit smoking.

Tom: Right, so let's break this down. Motivational Interviewing, or MI, is this really well-established counseling technique. It's not about telling someone to quit; it's about drawing out their own reasons for wanting to change.

Jane: Exactly. And the paper is from a team at the University of Toronto, including researchers from the Centre for Addiction and Mental Health. They're trying to solve a practical problem: how do you test whether an AI counselor is actually good at this?

Tom: And their answer is to build better fake patients. Instead of using real people, which is expensive and slow, they create AI-simulated clients that the AI counselors can practice on.

Jane: But here's the twist, Tom. The old way of doing this was kind of like giving an actor a full script and hoping they improvise well. The new way, which this paper calls Evoke-Sim, is more like giving the actor a character sheet and letting them reveal their backstory only when the counselor asks the right questions.

Tom: That's a great way to put it. The simulated clients have structured profiles, but they don't just spill everything at once. They have stages, and they only reveal information when the counselor earns it.

Jane: And that's what makes this evaluation so much smarter. It's not just checking if the counselor said nice things; it's checking if they said the right thing at the right time.

Tom: So we're talking about a higher bar for AI counselors, which is exactly what we need if these tools are going to be used by real people.

Jane: And that's the hook for the rest of our discussion. We're going to dig into how they built this, what they found, and why it matters for the future of AI in mental health.

Abstract: Tom: So Jane, we've set the stage. Now let's get into the meat of the abstract, because there's a lot packed into that summary.

Jane: There really is. The core problem they're tackling is that previous simulated clients weren't aligned with the actual tasks of Motivational Interviewing. MI has these four tasks: engaging, focusing, evoking, and planning.

Tom: And the one they focus on is evoking, which is all about drawing out the client's ambivalence and then strengthening their motivation to change. That's the heart of MI.

Jane: Right. So they built Evoke-Sim specifically for that evoking task. It's a multi-stage framework, which means the conversation flows through three distinct phases.

Tom: Stage one is evoking ambivalence, getting the client to talk about both the pros and cons of smoking. Stage two is evoking change talk, getting them to lean towards quitting. And stage three is evoking commitment to next steps.

Jane: And the key innovation here is the reveal policy. The simulated client has a full profile, but they only reveal motivations, barriers, and next steps when the counselor actively elicits them.

Tom: So if the counselor just starts asking about quitting plans in the first minute, the client will actually refuse to engage. That's a huge difference from previous systems where the client would just dump all their information.

Jane: And the results are striking. They showed that Evoke-Sim is much better at differentiating between good and bad counselors. A bad counselor, one that's confrontational and pushy, actually fails to get the client to progress through the stages.

Tom: Whereas a good counselor, one that uses proper MI techniques, gets the client to open up and move through all three stages successfully.

Jane: And importantly, they also showed that Evoke-Sim reduces non-grounded statements, meaning the simulated clients are less likely to invent facts that aren't in their profile. That's a big deal for realism.

Tom: So we're not just getting a better test; we're getting a more honest test. The simulated clients behave more like real people, which means the evaluation is more meaningful.

Jane: And that's the promise here. We're building a benchmark that actually tells us something about whether these AI counselors are ready for the real world.

Improvements: Tom: Jane, we've talked about what Evoke-Sim is. Now let's get into what it improves on, because the paper is very specific about the shortcomings of prior work.

Jane: And this is where it gets really interesting. The prior approaches, like the ones from Yosef and Wang, used structured profiles but let the simulated client access the full profile from the start.

Tom: Which sounds fine in theory, but in practice it creates a problem. The client becomes overly cooperative. They'll reveal motivations, barriers, and next steps all at once, even if the counselor didn't ask for them.

Jane: Exactly. And that makes the evaluation too easy. A bad counselor can stumble into success because the client is just handing over all the information.

Tom: So Evoke-Sim fixes that with the reveal policy we talked about. The client only reveals information when the counselor specifically asks for it, and even then, there are stage restrictions.

Jane: Right. In stages one and two, the client will only reveal preparatory change talk, things like desires, abilities, reasons, and needs. They won't talk about commitment or next steps until stage three.

Tom: And if the counselor tries to jump ahead, the client will refuse. They'll even terminate the conversation if the counselor keeps pushing for next steps too early.

Jane: That's a really important improvement. It means the evaluation is now testing not just whether the counselor uses MI skills, but whether they use them at the right time.

Tom: And the paper shows this works. Under the old Profile-Only system, even the bad counselor, MINA, achieved high completion rates. Under Evoke-Sim, MINA fails almost completely.

Jane: The numbers really tell the story. With Evoke-Sim, MINA only evokes ambivalence in forty-nine percent of conversations, and never reaches the end goal. But the good counselor, MI-Evoke, hits ninety-nine percent completion.

Tom: So the improvement here is about creating a test that actually separates the wheat from the chaff. It's not enough to sound empathetic; you have to guide the conversation in the right way.

Jane: And that's a massive step forward for evaluating AI counselors. We're moving from checking boxes to checking the quality of the therapeutic journey.

First Page: Tom: Jane, we've covered the big ideas. Now let's zoom in on the first page of the paper, because there's some context there that really frames the whole study.

Jane: The first page sets up the problem beautifully. It talks about how LLM-based MI counselors are becoming more common, and how people are actually using general-purpose chatbots for mental health counseling.

Tom: Which is a little scary when you think about it. People are turning to AI for serious mental health support, and we need to make sure those tools are actually good.

Jane: And the paper makes a really important point about evaluation. You can't just compare AI counselor utterances to human transcripts. That doesn't tell you if the counselor is actually effective.

Tom: Right, because a good MI counselor adapts to the client in front of them. You need real interactions with clients who are genuinely ambivalent about change.

Jane: And that's where the challenge comes in. Recruiting real clients is expensive, slow, and raises ethical concerns. You don't want to expose vulnerable people to an unqualified AI counselor.

Tom: So the solution is simulated clients. But the paper argues that previous simulations weren't good enough because they didn't align with the specific tasks of MI.

Jane: And that's the gap Evoke-Sim fills. It's not just a simulated client; it's a simulated client that understands the therapeutic process and responds accordingly.

Tom: The first page also introduces the team and their expertise. They're not just computer scientists; they're working with clinicians at the Centre for Addiction and Mental Health.

Jane: That collaboration is crucial. It means the stage progression and reveal policies aren't just arbitrary rules; they're grounded in real MI theory and practice.

Tom: And that gives the whole framework credibility. This isn't a toy; it's a serious tool designed with input from people who actually do this work.

Jane: So the first page really sets the stage for why this matters. We're not just building better AI; we're building safer, more effective AI for mental health.

Conclusion: Tom: Jane, we've covered a lot of ground on "Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients." Let's wrap this up.

Jane: Let's do it. The core takeaway is that Evoke-Sim provides a much more rigorous way to evaluate AI counselors by simulating clients that behave like real people with real ambivalence.

Tom: And the key innovation is the multi-stage progression and the reveal policy. The client doesn't just dump information; they reveal it only when the counselor earns it.

Jane: That means the evaluation tests not just whether the counselor uses MI skills, but whether they use them at the right time. And that's a huge step forward.

Tom: The results back it up. Under Evoke-Sim, a bad counselor fails almost completely, while a good counselor succeeds. The test actually separates the wheat from the chaff.

Jane: And they also showed that Evoke-Sim reduces non-grounded statements, so the simulated clients are more realistic and less likely to invent facts.

Tom: There are limitations, of course. The framework only covers the evoking task, not the other three MI tasks. And it's only tested in smoking cessation.

Jane: But that's not a criticism; that's a roadmap. Future work can extend this to other tasks and other behaviors. The foundation is solid.

Tom: And the implications are significant. As AI counselors become more common, we need rigorous ways to evaluate them. Evoke-Sim is a major step in that direction.

Jane: So we'll say goodbye to this paper, but we're excited to see where this line of research goes. Thanks for listening, everyone.

Tom: And as always, keep questioning, keep learning, and we'll see you on the next episode.

Jiading Zhu, Xinyu Cindy Wang, Thomas Nguyen, Yan Qing Lee, Osnat C. Melamed, Peter Selby, Jonathan Rose

University of Toronto · Centre for Addiction and Mental Health

cs.HC, cs.AI, cs.CL

Submitted: 2026-06-19

Updated: 2026-08-11

Comments: 53 pages

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 63/100

The gist: The paper presents Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating Motivational Interviewing (MI) counsellors in the context of smoking cessation,

Key concepts

Motivational Interviewing (MI)
MI is a counseling technique focused not on telling someone to change, but on drawing out their own reasons for wanting to change. It involves four tasks: engaging, focusing, evoking, and planning.
Evoke-Sim
This is the new framework for simulated clients. Instead of giving a client their full profile immediately, Evoke-Sim uses a multi-stage approach where the client only reveals motivations or next steps when the counselor asks specific questions.
Reveal Policy
This is a key innovation in Evoke-Sim. It dictates that the simulated client only reveals information—like barriers or plans—when the counselor actively elicits it, preventing clients from dumping all their information at once.
Task-Aware Multi-Stage Framework
This framework structures the conversation into three distinct phases: evoking ambivalence, evoking change talk, and evoking commitment. The simulated client's responses are restricted based on which stage of the MI process they are in.

Terminology

Summary

The paper presents Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating Motivational Interviewing (MI) counsellors in the context of smoking cessation, specifically designed for the Evoking MI task.

The paper addresses the challenge of evaluating LLM-based MI counsellors. MI is widely-used by human counsellors in domains such as smoking cessation and alcohol reduction. MI counsellors work through four tasks: Engaging, Focusing, Evoking, and Planning. The Evoking task is central because the counsellor elicits and acknowledges the client's ambivalence towards changing the unhealthy behaviour and gently guides the client towards change by eliciting the client's own motivations for change.

While prior work has achieved strong results in instructing LLMs to role-play as MI counsellors, proper evaluation requires interactions with clients ambivalent about change. Recruiting real clients is often long and costly, and raise practical and ethical concerns when participants are exposed to unqualified LLM-counsellors. Existing LLM-based simulated clients have not aligned with the specific tasks fundamental to the MI therapy approach.

The paper lists four main contributions:

  1. Evoke-Sim framework: a profile-grounded client simulation framework for task-aware, multi-stage evaluation of MI counsellors on the evoking MI task (code released).

  2. Task-aware evaluation metrics: a set of task-aware, stage-specific evaluation metrics across the three stages of evoking.

  3. Improved differentiation: We show that Evoke-Sim better differentiates MI counselling quality than generic profile-based simulated clients and prior metrics.

  4. Reduced non-grounded statements: We demonstrate that Evoke-Sim reduces non-grounded statements and premature disclosure relative to generic profile-based simulated clients.

  5. Released dataset: A released dataset of 83 diverse and structured profiles extracted from MI counselling sessions with human participants.

Evoke-Sim grounds each simulated client in a structured profile extracted from real MI counselling sessions between human smokers and an automated LLM-based counsellor. The profiles contain four fields:

  • Persona: general background of the client, including recent events, family relationships, occupation, religion and culture, health status, living situation, smoking history, and any other relevant background.

  • Motivations for Change: all the potential motivations for changing their smoking behaviour that the human client expressed.

  • Barriers Against Change: the client's expression of underlying barriers that would prevent them from changing their smoking behavior and maintain the status-quo.

  • Potential Next Steps: next steps the client is willing to consider or adopt to support their change and reduce smoking.

Motivations and barriers are labeled using DARNCATs-style coding grounded in MI theory, categorizing client language into change talk (Desire, Ability, Reason, Need, Commitment, Activation, Taking Steps, Other) and counter-change talk, with "+ denoting motivations and -" denoting barriers.

Profiles were extracted using an LLM-based profile-extraction pipeline from 106 eligible transcripts, with LLM-based filtering, manual review, and deterministic post-processing incorporating metadata (daily cigarette count, time to first cigarette, quit attempts, and Readiness Ruler scores for confidence, importance, and readiness). After processing, 83 distinct structured profiles were produced.

The evoking task is partitioned into three connected stages, defined in consultation with MI experts:

  • Stage 1 (Evoking Ambivalence): Goal is to evoke both sides of ambivalence, defined as evoking at least one motivation item and one barrier item. Transition to Stage 2 occurs when this is achieved. Early termination occurs if transition requirements are not met within 10 client turns; if next steps or commitment were asked 3 times.

  • Stage 2 (Evoking Change Talk): Goal is to evoke change talk and reduce sustain talk. A sliding window change score over a 10-turn window is maintained, where the score increases with each change talk utterance (labelled +) and decreases with each sustain talk utterance (-). Stage 2 transitions to Stage 3 when the change score reaches at least 3. Early termination occurs if the change score reaches the floor; if next steps or commitment were asked 3 times.

  • Stage 3 (Evoking Commitment to Next Steps): Goal is to evoke potential next steps, then commitment. The end goal is satisfied when Commitment+ is evoked after any potential next steps have been discussed. Early termination occurs if the change score reaches the floor.

Progression is strictly sequential: the client simulation starts in Stage 1 and can only move forward one stage at a time (Stage 1 → Stage 2 → Stage 3), with no stage skipping or backward transitions.

The change score floor is unique to each client profile, initialized from the three Readiness Ruler scores (confidence, importance, readiness), computed as smin = -3 - k where k = clip((R-4)/5, 0, 3) and R = c + i + r, giving a range of smin ∈ [-6, -3].

A key feature is that information disclosure from the simulated client is gated turn by turn, so that only appropriate profile items are revealed. At the start, only the persona field items from the structured profile are made visible, while motivations, barriers, and next steps are hidden. Throughout the conversation, revealed items remain cumulatively visible.

The reveal controller makes one decision per turn: reveal at most one new item, make no new reveal for that turn, or instruct a refusal. Items are only revealed if the counsellor specifically asks about them. When 3 or more types of items are requested at the same turn, no new item is revealed. In Stages 1 and 2, only motivations and barriers in preparatory DARNCATs categories (Desire, Ability, Reason, and Need) can be revealed, while requests for mobilizing categories (Activation, Commitment, Taking Steps) and next steps are refused. In Stage 3, all profile items are allowed to be revealed. When counsellor behaviour is not adherent to MI principles, the controller forces either a refusal, or a sustain response grounded in visible barriers.

Evoke-Sim consists of seven internal modules: Counsellor Intent Classifier, Counsellor Behaviour Classifier, Client Behaviour Classifier, Conversation State Controller, Reveal Controller, Prompt Builder, and Client Utterance Generator. The three classifiers and the Client Utterance Generator are LLM-based; the two controllers and the Prompt Builder are rule-based.

The experiments use two client frameworks (Evoke-Sim and Profile-Only baseline) and three MI counsellors, all using OpenAI gpt-5.2-2025-12-11:

  • MINA (MI Non-Adherent counsellor): a baseline designed to deliberately not adhere to MI principles, prompted to be persuasive and confrontational about change.

  • MIBot v6.3A: a MI-adherent counsellor from a prior study, emphasizing MI skills such as reflective listening and eliciting ambivalence.

  • MI-Evoke: a counsellor designed specifically for the evoking task, using the same base prompt as MIBot v6.3A but with strategy-aligned prompting from a constrained MI strategy decision space.

The full factorial matrix of 2 client frameworks × 3 counsellors results in 6 experimental arms, all evaluated on the same 83 profiles.

Under Profile-Only, there is not a clear separation in performance between the counsellors across all task-aware metrics. All three counsellors achieve high task-outcome scores, including MINA. End goal completion rates are 0.89 (MINA), 0.94 (MIBot v6.3A), and 1.00 (MI-Evoke).

Under Evoke-Sim, "the same counsellors are clearly separated by the same metrics. MINA fails to reach the end goal under Evoke-Sim, with less than half of conversations successfully evoking ambivalence. MI-Evoke reaches near-ceiling end goal completion with strongly positive Final CS. MIBot v6.3A remains substantially better than MINA, but remains far below MI-Evoke on end goal completion and stage progression."

For MITI metrics, MITI metrics mainly show a large gap between MINA and the other counsellors, while providing limited separation between MIBot v6.3A and MI-Evoke under Profile-Only. Under Evoke-Sim, the top global MITI performance shifts to MI-Evoke (Global, CCT, SST, and Partnership), while MIBot v6.3A stays close and higher on Empathy and on the summary measures (%CR and R:Q).

Mean non-grounded rates are 0.070 for human clients, 0.55 for Profile-Only, and 0.47 for Evoke-Sim with MIBot v6.3A as counsellor. Relative to Profile-Only, Evoke-Sim reduces client-side non-grounded statements under matched counsellor conditions, while still remaining above the real-transcript baseline.

For information disclosure, "Evoke-Sim clients generally reveal non-persona information more slowly and to a lower final level than Profile-Only clients. Notably, Profile-Only conversations tend to share a large fraction of non-persona profile items within the first few turns (as many as 20%), which is much higher than real transcripts, while Evoke-Sim conversations show more gradual and controlled information disclosure over turns."

"We propose Evoke-Sim, a task-aware multi-stage client simulation framework for evaluating MI counsellors on the evoking MI task. Results show that Evoke-Sim provides stronger task-aware differentiation of counsellor quality than existing profile-grounded clients, while also making MITI global scores more aligned with task-specific performance. Additionally, Evoke-Sim is shown to reduce client-side non-grounded statements and slow information disclosure. Overall, Evoke-Sim sets a higher standard for LLM-based MI counsellors by holding them to task-specific performance in evoking."

The paper acknowledges several limitations: the framework is scoped... limited to only the 'Evoking' MI task; source transcripts are full MI sessions that can extend beyond evoking, creating a potential scope mismatch; the extracted profile set may be incomplete relative to real underlying client profiles; experiments are only in the smoking-cessation context, and in a text-only interface; the three-stage partition was developed through collaboration with MI experts, but alternative expert interpretations are possible; LLM-based modules can still introduce inaccuracies and do not replace expert human coding; and this work evaluates counsellor quality in simulation rather than direct patient outcomes.

The paper notes that Evoke-Sim could be misused to generate realistic synthetic counselling clients outside of research settings and could encapsulate narrow assumptions about what counts as high quality MI counselling. It positions Evoke-Sim as a research benchmark rather than a deployment tool. All data comes from a publicly available dataset released by a prior study in which participants consented to use of their data for research purposes, distributed under CC BY-SA 4.0, with data anonymized and manually checked for personally identifying information and offensive content.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:


Improvement:

Implement a client simulation module that partitions the Motivational Interviewing (MI) Evoking task into three sequential stages—(1) Evoking Ambivalence, (2) Evoking Change Talk, (3) Evoking Commitment to Next Steps—with deterministic transition rules and early-termination conditions.

What the improved AI system can do:

  • Simulate a client that progresses only when the counsellor successfully elicits required client language (e.g., at least one motivation and one barrier in Stage 1).

  • Automatically terminate the conversation if the counsellor asks for next steps or commitment more than three times before Stage 3, or if the client’s change score falls below a profile-specific floor.

  • Prevent stage-skipping and backward transitions, ensuring the evaluation reflects true MI task progression.

The improved AI system can:

  • Evaluate MI counsellors with task-aware, stage-specific metrics that clearly separate high-quality from low-quality counselling.

  • Simulate realistic, ambivalent clients that disclose information only when appropriately elicited, resist premature planning, and terminate unproductive sessions.

  • Reduce hallucination and premature disclosure in simulated clients, making evaluations more trustworthy.

  • Align MITI global scores with actual task performance, providing a more holistic and accurate assessment.

  • Provide a reusable, open-source benchmark for the MI evoking task, accelerating research and development of automated counselling systems.

Abstract

The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundamental to the MI therapy approach. A key task is evoking, in which the counsellor first elicits the client's ambivalence and then strengthens the client's motivation for change. We present Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating MI counsellors in smoking cessation, designed specifically for the evoking MI task. Evoke-Sim employs structured client profiles, an evoking-specific three-stage conversation flow, and a reveal policy that regulates which client profile information might be disclosed at each stage. We show that compared to existing profile-grounded simulated clients, Evoke-Sim is better at differentiating levels of MI quality using task-aware evaluation metrics, while reducing non-grounded client statements and premature disclosure of client information, setting a higher standard for the evaluation of LLM-based MI counsellors.

Sources

Related papers