Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Reichman University
cs.CL, cs.AI
Submitted: 2026-08-12
Updated: 2026-09-01
Comments: 13 pages, 4 figures. Accepted at an ICLR 2025 workshop
Code: https://github.com/YuxiXie/MCTS-DPO
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
Terminology
Summary
Published: Conference paper at ICLR 2025
The paper introduces a novel framework called Preference Tree Optimization (PTO) designed to iteratively improve agent models in goal-oriented dialogue systems. The framework generates preference data using a method called Preference Tree with Look-Ahead. The research focuses on Motivational Interviewing (MI) — a counseling technique aimed at facilitating behavioral change — leveraging virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets. By combining this method with Direct Preference Optimization (DPO), the authors aim to enhance the agent's decision-making capabilities over iterative training cycles. The framework addresses data scarcity and advances the development of more nuanced and effective dialogue systems in goal-oriented domains.
Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing (MI). Models trained with PTO consistently outperformed the baseline in key metrics such as session satisfaction and working alliance. Additionally, incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies, with deeper look-ahead configurations yielding the most stable and high-scoring results.
Goal-oriented dialogue systems are designed to achieve specific objectives through interactive conversations. Developing such systems in specialized domains is challenging due to the complexity of interactions and the scarcity of domain-specific data. Motivational Interviewing (MI) is such a domain — a counseling approach that facilitates behavioral change through collaborative, client-centered dialogue, requiring nuanced understanding and adaptability from the conversational agent.
The research introduces a framework for iteratively improving agent models in goal-oriented dialogue systems, called Preference Tree Optimization (PTO), by generating preference data using a novel method called Preference Tree with Look-Ahead. This method systematically simulates various conversational paths and evaluates them using an oracle to generate preference data. This preference data is used with Direct Preference Optimization (DPO) to iteratively refine the agent model, enhancing its decision-making capabilities.
The approach leverages existing virtual patients and evaluators from previous research in MI, making it an ideal testbed for the framework. Similar preference-based strategies have improved models in well-defined analytic tasks like games, coding, and math, but their application to human-centric domains like Motivational Interviewing — where objectives are subjective and nuanced communication is key — remains largely unexplored.
The PTO framework operates in two iterative steps:
-
Preference Data Generation: The User Model is prompted with a range of attributes to simulate diverse user personalities. For each digital user personality, the Preference Tree with Look-Ahead method is used in conjunction with the Oracle Evaluator and the current agent model to generate a preference tree that explores various conversational pathways. These trees are aggregated into a comprehensive preference dataset.
-
Model Training: The current agent model is trained on the newly generated preference dataset using Direct Preference Optimization (DPO), resulting in an improved model. The updated agent model is then used for the next iteration, repeating the process for continuous improvement.
The PTO framework is designed exclusively as an offline training paradigm. Although the DPO process is computationally intensive — since it is applied at each simulated decision point during training — this cost is incurred only once during model development. Once trained, the automated therapist is deployed for real-time conversation, where the inference process is fast and efficient.
The work makes several contributions:
-
Introduction of the Preference Tree with Look-Ahead, a novel method that systematically simulates and evaluates potential conversational trajectories to generate high-quality preference data.
-
Proposal of the Preference Tree Optimization (PTO) framework, which integrates preference data with DPO to iteratively refine an agent's decision-making capabilities over successive training cycles.
-
Validation of the approach in the challenging domain of Motivational Interviewing (MI) by leveraging virtual patients and oracle evaluators to simulate realistic, high-stakes conversational scenarios.
-
The methodology offers broad insights and generalizable strategies for applying preference-based optimization to other specialized dialogue domains.
Recent breakthroughs in NLP and Large Language Models (LLMs) have dramatically advanced dialogue systems. However, designing goal-oriented systems for specialized domains like MI remains challenging due to limited availability of domain-specific data and the complexity of managing nuanced, multi-turn interactions. Pure generative models may not naturally exhibit goal-directed behavior. While reinforcement learning (RL) offers a potential path to integrating goal orientation, identifying suitable reward functions in domains like psychology is far from straightforward.
Traditional approaches like Reinforcement Learning from Human Feedback (RLHF) involve training a separate reward model based on human evaluations of model outputs, which then guides the language model through reinforcement learning. Although effective, RLHF can be complex and resource-intensive.
Direct Preference Optimization (DPO) offers a more streamlined alternative by directly optimizing the language model using preference data, eliminating the need for a separate reward model and the complexities of reinforcement learning. DPO establishes a direct mapping between LLM policies and reward functions.
The paper categorizes related approaches into three groups:
-
Score-Based Synthetic Data Generation: Methods like West-of-N (Pace et al., 2024) leverage language models to produce multiple candidate responses, with a reward model scoring them to form synthetic preference pairs. Online AI Feedback (OAIF) (Guo et al., 2024) employs an LLM as an annotator to provide on-the-fly feedback on pairs of responses sampled from the current model.
-
Self-Evaluation–Driven Improvement: Methods like Self-Rewarding Language Models (Yuan et al., 2024b) generate multiple responses and use an LLM-as-a-judge to rank them. I-SHEEP (Liang et al., 2024) synthesizes data, self-assesses quality, and filters out low-quality responses before applying supervised fine-tuning.
-
Search-Based and Tree-Structured Approaches: Methods like Monte Carlo Tree Search (MCTS) integrated with iterative preference learning (Xie et al., 2024), Preference Trees (Yuan et al., 2024a), and prompt-based search methods (Yu et al., 2023).
The paper notes: "Our work bridges search-based and score-based paradigms: the Preference Tree with Look-Ahead method employs tree-structured exploration of conversational trajectories, while the oracle evaluator provides score-driven comparisons, enabling iterative refinement via Direct Preference Optimization (DPO) to enhance goal-oriented dialogue agents in specialized domains such as Motivational Interviewing. Unlike prior work, which has primarily focused on structured tasks such as coding, math, or games, our approach explores preference-based optimization in a domain that requires deep human understanding, where objectives are inherently subjective and harder to quantify."
Motivational Interviewing (MI) is a client-centered counseling approach aimed at eliciting behavioral change by helping clients explore and resolve ambivalence. Implementing MI in AI dialogue systems presents unique challenges due to the need for empathy, adaptability, and the ability to interpret subtle conversational cues.
Previous research (Yosef et al., 2024) utilized AI-generated patient simulations to assess MI sessions, demonstrating the feasibility of virtual patients in training and evaluating therapeutic dialogues. Unlike methods that rely on pre-existing datasets, the PTO framework iteratively generates training data from simulated conversations using the Preference Tree with Look-Ahead method, refining the model at each iteration via DPO.
The Preference Tree with Look-Ahead method systematically explores potential conversational paths by simulating multiple agent responses and their subsequent dialogue trajectories. The process is as follows:
-
Agent Decision Point: At each turn, the agent model generates N possible responses.
-
Branch Initialization: For each response, a new branch is created, and the response is appended to the conversation history.
-
Look-Ahead Simulation: Each branch simulates K future steps, alternating between the agent and the virtual patient, to anticipate the long-term implications of the agent's response.
-
Oracle Evaluation: An oracle evaluator assesses each branch based on predefined criteria (e.g., adherence to MI principles, empathy, goal progression) and assigns scores.
-
Preference Recording: The response with the highest score is considered the preferred response, and the one with the lowest score is the least preferred. The preference tuple is recorded in the dataset.
-
Conversation Update: The conversation continues with the preferred response, and the process repeats until a termination condition is met (e.g., reaching maximum conversation length or achieving the goal).
By considering future conversation trajectories, the agent is expected to learn to make decisions that are not only immediately appropriate but also beneficial in the long term.
The agent model is iteratively improved through cycles of preference data generation and training using DPO:
-
Initial Training: The agent model is initially trained on available data or pre-trained weights.
-
Preference Data Generation: Using the current agent model, the Preference Tree with Look-Ahead method generates new preference data, capturing the agent's strengths and weaknesses.
-
Preference Data Filtering: A preference sample is retained only if the winning score surpasses the losing score by a predefined threshold (0.1 in experiments), ensuring only clearly distinguishable preference pairs contribute to training.
-
Model Update: The agent model is fine-tuned using DPO on the newly generated preference data.
-
Evaluation: The updated model is evaluated using predefined metrics.
-
Iteration: Steps 2-5 are repeated, allowing the agent to improve over time through continuous learning.
This process balances exploration (generating new conversational paths) and exploitation (refining the agent's responses), leading to incremental enhancements in performance.
-
Agent Model: Llama-2-7B was used as the base model for the therapist agent.
-
User Model: Virtual patients were simulated using GPT-3.5, based on guidelines from previous MI research. Each patient is defined by parameters such as gender, age, problem (smoking/obesity), duration, prior attempts to resolve the issue, and cooperation level, creating 96 unique profiles to capture diverse challenges and attitudes toward counseling.
-
Oracle Evaluator: GPT-3.5 model was used as the oracle evaluator, using specific questionnaires designed to assess MI adherence and conversational quality. The final score is calculated as the average of the two questionnaire scores.
Notably, GPT-3.5 was used in fixed, separate roles (with distinct prompts for user simulation and oracle evaluation), and it was not updated or fine-tuned at any point during the training process.
-
Look-Ahead Depths: Two different look-ahead depths were tested: 0 (no look-ahead) and 5.
-
Iterations per Look-Ahead: For each look-ahead depth, 7 iterative training cycles were conducted.
After each iteration, 96 separate conversations with virtual patients were conducted to evaluate the agent's performance. Each conversation was scored by the oracle evaluator based on two distinct questionnaires measuring MI adherence and overall conversational quality.
The agent's effectiveness was evaluated using two primary metrics:
-
Session Satisfaction (Q1): Aggregates scores assessing overall satisfaction, content relevance, motivation facilitation, learning outcomes, and applicability to everyday life.
-
Working Alliance (Q2): Aggregates scores evaluating the therapist's interpersonal skills, empathy, communication effectiveness, and ability to establish a collaborative relationship.
-
Final Score: Calculated as the average of Session Satisfaction and Working Alliance scores.
Table 1: Average Performance Scores and Standard Deviations Across Models
Model Session Satisfaction (Q1) Mean SD Working Alliance (Q2) Mean SD Final Score Mean SD
Base 3.521 1.056 3.385 0.539 3.453 0.740
L0 M1 3.863 1.012 3.452 0.731 3.657 0.824
L0 M2 3.750 1.059 3.435 0.788 3.593 0.878
L0 M3 3.796 0.868 3.567 0.511 3.682 0.649
L0 M4 3.969 0.979 3.585 0.642 3.777 0.769
L0 M5 3.744 1.124 3.478 0.687 3.611 0.856
L0 M6 3.794 1.143 3.494 0.633 3.644 0.834
L0 M7 3.677 1.098 3.452 0.667 3.565 0.828
L5 M1 3.898 1.005 3.523 0.480 3.710 0.712
L5 M2 3.969 0.809 3.618 0.455 3.794 0.594
L5 M3 4.050 0.818 3.683 0.548 3.866 0.611
L5 M4 3.981 0.801 3.605 0.351 3.793 0.524
L5 M5 4.225 0.775 3.660 0.451 3.942 0.559
L5 M6 4.112 0.868 3.656 0.477 3.884 0.629
L5 M7 4.190 0.614 3.775 0.332 3.982 0.414
Key findings from the results:
-
All PTO-trained models outperform the baseline across all evaluated metrics, demonstrating that preference-based optimization improves goal-oriented dialogue performance.
-
Models trained with deeper look-ahead (depth-5) achieve higher scores than those trained with no look-ahead (depth-0), suggesting that anticipating future conversational paths enhances both session satisfaction and the working alliance.
-
L5 M7 exhibits the lowest variance across Q1, Q2, and Final Score (underlined in the table), suggesting that deeper look-ahead not only enhances performance but also ensures more consistent and reliable motivation interventions.
-
PTO-trained models tend to reduce conversation length compared to the baseline. L5 M7 achieves the most substantial reduction, decreasing the average number of dialogue turns from 43.7 (baseline) to 34.4.
A one-way ANOVA confirms that model choice significantly influences Q1, Q2, Final Score, and conversation length:
Table 2: One-Way ANOVA Results
Metric F-Statistic p-value
Final Score 15.637 3.60e-07
Session Satisfaction (Q1) 13.654 2.17e-06
Working Alliance (Q2) 13.446 2.63e-06
Conversation Length 11.928 1.06e-05
Post-hoc Tukey HSD tests compared the baseline model against the best-performing models from each look-ahead depth: L0 M4 (best depth-0 model) and L5 M7 (best depth-5 model). Results indicate that both L0 M4 and L5 M7 significantly outperform the baseline across all three metrics (Q1, Q2, and Final Score). While L5 M7 achieves the highest Final Score, its improvement over L0 M4 is only statistically significant for Q2, indicating that deeper look-ahead particularly strengthens the working alliance.
The experimental results demonstrate that the PTO framework consistently improves dialogue performance compared to the baseline. Importantly, these improvements were achieved using a base pre-trained model (Llama-2-7B) that was neither instruction-tuned nor fine-tuned via supervised learning; instead, all training was conducted solely with data generated by the Preference Tree with Look-Ahead method.
Both look-ahead configurations (depth-0 and depth-5) yield significant gains in Session Satisfaction (Q1), Working Alliance (Q2), and overall Final Score. Notably, the best-performing depth-5 model (L5 M7) not only achieved the highest scores but also exhibited the lowest variance, indicating more stable and reliable interactions. This suggests that incorporating look-ahead enables the agent to anticipate future conversational turns, leading to more effective, empathetic, and streamlined dialogues.
The paper acknowledges potential biases in automated evaluation:
-
Positional bias: Minor variability in evaluation of initial utterances, but the oracle's scoring remains largely consistent throughout the dialogue.
-
Preference bias: The evaluator might favor certain stylistic or content-related features, which could lead the agent to optimize for superficial attributes rather than genuine conversational quality — a type of
reward hacking.
The paper notes that both the oracle evaluator and virtual patients are implemented as fixed, pre-trained models, so using the same underlying model for both roles is not the primary source of risk for reward hacking. Importantly, the oracle evaluator was validated by human assessments — although the correlation was moderate, this validation indicates that the evaluation criteria capture meaningful aspects of effective counseling.
Future work will focus on:
-
Elucidating whether and why deeper look-ahead (L5) offers advantages over no look-ahead (L0) in this soft domain.
-
Investigating underlying mechanisms that contribute to improved long-term planning, such as better management of conversational dynamics and enhanced anticipatory decision-making.
-
Benchmarking against leading state-of-the-art methods — specifically the online alignment framework from Guo et al. (2024) and the self-rewarding language model approach from Yuan et al. (2024b).
This work was partially supported by the European Commission Horizon 2020 project GuestXR (#101017884).
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
-
Improvement: Integrate a tree-structured search mechanism that simulates multiple future conversational trajectories (K steps ahead) before selecting a response, rather than making greedy, turn-by-turn decisions.
-
Capability: The AI can now evaluate the downstream consequences of its current utterance, choosing responses that optimize long-term outcomes (e.g., building rapport, achieving a counseling goal) instead of only immediate coherence.
-
Improvement: Build a closed-loop training pipeline where the AI generates its own preference pairs from simulated interactions, filters them by a score threshold (e.g., 0.1), and fine-tunes itself via Direct Preference Optimization (DPO) over multiple cycles (e.g., 7 iterations).
-
Capability: The AI continuously improves without requiring new human-annotated data, adapting to edge cases and diverse user profiles (e.g., 96 distinct virtual patient personalities) that it encounters during self-play.
-
Improvement: Use a separate, fixed evaluator model that scores each simulated branch on multiple criteria (e.g., session satisfaction, working alliance) and averages them into a final score, guiding preference selection.
-
Capability: The AI learns to balance competing objectives (e.g., empathy vs. goal progression) rather than optimizing a single metric, leading to more holistic performance in subjective domains.
-
Improvement: Expose look-ahead depth (e.g., 0 vs. 5) as a configurable hyperparameter, with deeper look-ahead shown to reduce variance in outcomes (e.g., standard deviation dropped from 0.740 to 0.414 in final scores).
-
Capability: The AI can be tuned to produce more predictable, reliable interactions in high-stakes settings (e.g., therapy), where consistency is as important as peak performance.
-
Improvement: Retain only preference pairs where the winning score exceeds the losing score by a predefined threshold, discarding ambiguous or borderline examples.
-
Capability: The AI avoids learning from noisy or low-confidence signals, resulting in more robust policy updates and faster convergence during iterative training.
-
Improvement: Start with a pre-trained model (e.g., Llama-2-7B) that is not instruction-tuned, and train it solely on PTO-generated preference data.
-
Capability: The AI can be adapted to specialized domains (e.g., motivational interviewing) without requiring expensive supervised fine-tuning datasets, making the approach accessible for resource-constrained scenarios.
-
Anticipate user responses: Simulate 5 future turns to predict whether a question will lead to engagement or resistance, then adjust its phrasing accordingly.
-
Maintain long-term rapport: Choose responses that strengthen the working alliance over the entire session, not just the current exchange, leading to higher user satisfaction scores (e.g., 3.982 vs. 3.453 baseline).
-
Reduce conversation length: Achieve goals in fewer turns (e.g., 34.4 vs. 43.7 turns), making interactions more efficient without sacrificing quality.
-
Handle diverse user profiles: Adapt to 96+ distinct user personalities (e.g., varying cooperation levels, problem durations) by learning from simulated interactions with each profile.
-
Self-improve without human labels: Generate its own training data through simulated interactions, filter for quality, and iteratively refine its policy—useful for domains where expert annotations are scarce.
-
Optimize for subjective outcomes: Learn to balance multiple, hard-to-quantify objectives (e.g., empathy, clarity, goal attainment) by using an oracle that scores along multiple dimensions.
-
Provide stable performance: Reduce variance in output quality, making the system more reliable for deployment in sensitive applications (e.g., mental health, education).
-
Transfer to other soft domains: Apply the same framework to negotiation, tutoring, or customer service, where long-term strategy and nuanced communication are critical.
Sources
- Deep reinforcement learning from human preferences
- Direct Language Model Alignment from Online AI Feedback
- I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
- Training language models to follow instructions with human feedback
- West-of-N: Synthetic Preferences for Self-Improving Reward Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
- Prompt-Based Monte-Carlo Tree Search for Goal-Oriented Dialogue Policy Planning
- Advancing LLM Reasoning Generalists with Preference Trees
- Self-Rewarding Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering