Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

arXiv:2608.10626 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Yi Wei, Shuo Jiang, Huaixia Dou, Jie Zhu, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang

Alibaba Cloud Computing · Beihang University · Soochow University

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 10 pages, 4 figures, 6 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: The paper introduces a dual-loop self-evolution framework for multi-turn empathetic dialogue, driven by verifiable emotion feedback.

Terminology

Summary

The paper introduces a dual-loop self-evolution framework for multi-turn empathetic dialogue, driven by verifiable emotion feedback. The core problem identified is that existing reinforcement learning (RL) methods evolve the dialogue policy while keeping the training interaction distribution fixed, creating a mismatch between policy competence and training experience. The authors state: "existing methods still treat the training interaction distribution as a predefined and fixed external condition... Emotion rewards continually update the dialogue policy, but do not change how subsequent experience is allocated."

The framework consists of two nested loops: "the inner loop uses continuous emotion rewards to improve how the policy responds, while the outer loop reuses group outcomes to identify what the policy should practice next and reallocates subsequent interactions accordingly." The user simulator and verifier remain frozen throughout training. The interaction space is defined by three axes—disclosure readiness (3 levels), emotional activation (2 levels), and relational trust (4 levels)—yielding 24 interaction states. These states are combined with support intents to form 192 controller units.

For policy learning, the framework uses group-relative policy optimization (GRPO) with continuous emotion rewards. For the outer loop, it computes thresholded group pass rates, then estimates policy-relative interaction utility using a leave-one-intent prior for hierarchical evidence sharing, a boundary score that maximizes when success and failure coexist, and an uncertainty bonus for under-observed units. The allocation formula is: pt(zc) = (1−ϵ) A c,z 1/γ / Σ z'∈Z A c,z'1/γ + ϵ/Z, where γ is sampling temperature and ϵ reserves uniform rehearsal.

Key experimental results on the SAGE benchmark show the framework raises Qwen3-8B Overall from 53.87 to 79.24 and outperforms protocol-matched uniform emotion-reward reinforcement learning by 7.23 points. Across three independent training runs, the weakest complete-framework run (78.11) still exceeds the strongest uniform RL run (73.76). The framework also improves ESConv metrics (Distinct-2 from 28.18 to 33.58), achieves the highest ESC-Eval aggregate score (2.55), improves EIBench from −9.79 to −7.55, and raises human overall quality from 3.0 to 3.5.

Ablations show that removing adaptive allocation (Interaction State Only) drops SAGE to 69.42, scenario-only allocation drops to 70.03, removing hierarchical sharing costs 3.04 points, removing the uncertainty bonus costs 4.19 points, and the variance-gated controller causes KL instability with a score of 61.45. The interaction-state space is validated behaviorally: a held-out judge recovers disclosure, activation, and trust levels with accuracies of 86.7%, 92.5%, and 82.5%, respectively, while 92.5% of trajectories retain the original persona, event, and hidden need. The framework requires no additional rollouts, model passes, or simulator/verifier calls—only constant-time statistic updates.

Improvements for AI systems

Improvements to AI systems:

  1. Adaptive training-distribution allocation via outer-loop feedback – Replace fixed or uniformly sampled training distributions in RL fine-tuning with a dynamic allocation mechanism that reweights interaction states based on policy-relative utility (boundary score + uncertainty bonus + hierarchical evidence sharing). The improved AI system can automatically shift practice toward underperforming or under-observed interaction types, reducing competence–experience mismatch and accelerating convergence without extra rollouts.

  2. Continuous emotion-reward-driven policy optimization (GRPO with verifiable emotion signals) – Integrate fine-grained, verifiable emotion feedback (e.g., disclosure readiness, emotional activation, relational trust) as continuous rewards into group-relative policy optimization. The improved AI system can learn nuanced empathetic response policies that maximize emotional alignment, not just task success, leading to higher human-perceived quality (e.g., +0.5 on human overall quality).

  3. Variance-gated controller with KL-stability safeguard – Replace unstable variance-gated controllers with a temperature-sampled allocation formula (γ) plus uniform rehearsal (ϵ) to maintain KL stability. The improved AI system can explore diverse interaction states without policy collapse or divergence, preserving training robustness across multiple runs.

  4. Hierarchical evidence sharing via leave-one-intent prior – Use a leave-one-intent prior to share statistical evidence across related interaction states (e.g., different support intents within the same scenario). The improved AI system can make reliable allocation decisions even with sparse data, improving sample efficiency and generalization to unseen interaction combinations.

  5. Uncertainty-bonused allocation for under-observed units – Add an uncertainty bonus to the interaction utility estimate, prioritizing states with low observation counts. The improved AI system can actively seek out rare but critical interaction scenarios (e.g., high-trust, high-disclosure, low-activation states), preventing neglect of edge cases and improving robustness in real-world deployment.

  6. Frozen simulator/verifier with constant-time statistic updates – Keep user simulator and verifier frozen while only updating allocation statistics in O(1) time. The improved AI system can perform self-evolution with negligible computational overhead, enabling continuous on-device or low-resource adaptation without retraining or additional inference calls.

  7. Behaviorally validated interaction-state space – Use a 24-state interaction space (disclosure × activation × trust) that is recoverable by a held-out judge (86.7–92.5% accuracy). The improved AI system can reliably classify and track user emotional state in real time, enabling context-aware empathetic responses that preserve persona and hidden needs (92.5% trajectory fidelity).

What the improved AI system can do:

  • Self-evolve its dialogue policy by dynamically reallocating training interactions to the most informative emotional states, achieving +25.37 points on SAGE Overall (53.87→79.24) and outperforming uniform RL by +7.23 points.

  • Maintain stable training across multiple runs (weakest run 78.11 > strongest uniform RL 73.76), avoiding KL divergence and policy collapse.

  • Improve response diversity and empathy on benchmarks (Distinct-2 +5.40, ESC-Eval 2.55, EIBench −9.79→−7.55) without extra rollouts or simulator calls.

  • Operate in resource-constrained environments due to constant-time statistic updates, enabling lifelong learning in production chatbots or embodied agents.

  • Adapt to sparse or novel interaction contexts via hierarchical sharing and uncertainty bonuses, making it suitable for domains with limited user interaction data.

Sources

Related papers