EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation".
Jane: The paper was written by Yunbo Long, Haolang Zhao, Lukas Beckenbauer, Liming Xu and Alexandra Brintrup from University of Cambridge and Technical University of Munich and Exiger LLC and The Alan Turing Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, how does EmoDistill actually work? It sounds like a complex pipeline to build these agents.
Jane: Essentially, the authors are using a clever distillation process to teach emotional skills without running millions of costly live negotiations, which is a huge operational win.
Lu: They define an "emotional negotiation skill" as something that binds the dialogue state and its emotional stance—it’s a reward-annotated turn from an offline LLM-vs-LLM sweep.
Meng: This entire process relies on three stages: first, using an Implicit Q-Learning selector to choose the emotion, then Supervised Fine-Tuning via LoRA to learn how the express it in a small model, and finally refining it with Judge Policy Optimization or JPO.
Jane: The authors are decoupling the *selection* of emotion from its *expression*, so they aren't just telling an LLM to "be angry"; they are training a smaller, targeted SLM to execute that specific skill.
Tom: It’s like creating a specialized agent that has been trained by watching and learning from top performers, which is incredibly powerful.
Lu: And Lalam points out that this system learns not just *what* to say but *when* to say it, ensuring the emotional framing is perfectly timed for maximum leverage.
Meng: The engineering benefit of this specific mechanism is that we are taking a large, complex skill and distilling it into a manageable student model rather than relying on brittle prompt templates.
Lalam: It allows us to build agents that possess a form of genuine strategic competence by using structured knowledge instead of hoping they just happen to be the right emotion.
Improvements: Tom: Moving from *how* it works to *what* it achieves, the results are genuinely impressive across all four domains.
Jane: The full EmoDistill policy consistently outperforms vanilla LLM and SLM baselines, which is not just a slight improvement but a substantial difference in success rate and utility.
Lu: What’s particularly interesting is that the gains aren't just from the high-level emotion selection; they are proving that optimizing *how* the expression—the utterance itself—is realized is crucial for achieving better outcomes.
Meng: Looking at Table one it clearly shows that combining those three components, IQL, SFT, and JPO, provides a synergistic effect that simply choosing the right emotion alone cannot achieve.
Tom: And it’s not just in one specific area; the fact that EmoDistillC trained on Credit Recovery performs well when applied to Student Sleep Scheduling demonstrates cross-domain transferability.
Jane: It does, and this confirms what the paper' suggests: emotional cues are not just a cosmetic style but a core component of the strategy itself.
Lu: We can also see some fascinating trade-offs in their findings; for example, IQL+SFT+JPO often takes more rounds than other methods, but that is because they are achieving significantly better value per turn.
Meng: The engineering takeaway here is that the system isn't just succeeding; it’s succeeding in a way that is strategically superior to simply rushing an agreement.
Lalam: This proves that we can build reliable, high-performing agents by learning complex emotional dynamics instead of relying on simple, brittle heuristics.
Conclusion: Tom: We've seen how EmoDistill works and what it achieves—it’s a powerful tool for teaching strategic negotiation in AI agents. But before we wrap up, let's hear the final thoughts from everyone.
Jane: I hope the authors’ findings are encouraging; if an AI can learn to be strategically effective through emotional framing, it has real potential for making complex negotiations much more efficient across industries.
Lu: I'm extremely excited about seeing these skills transferred to even more advanced applications, beyond just simple agent-to-agent scenarios; the scope of this work is massive.
Meng: My main hope is that this offline approach provides a stable foundation for building robust commercial AI tools that can handle complexity without the instability we see in current online reinforcement learning methods.
Lalam: I think the ultimate impact will be in how this changes our perception of what a truly competent autonomous agent is, moving beyond just efficiency to including strategic emotional intelligence.
Tom: These are all incredibly powerful visions; we appreciate you sharing your insights today on "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation."
Final Goodbye: Jane: We've seen a framework that teaches AI to negotiate effectively, not just by chance, but through learned emotional strategy.
Lu: I agree; the fact they can use this offline training to unlock future possibilities is truly thrilling.
Meng: The ability EmoDistill provides for avoiding those costly live negotiations is what makes this approach scalable and practical.
Lalam: It allows us to build agents that feel and operate with genuine strategic depth, which will improve our relationship with AI assistants moving forward.
Tom: It's a framework that teaches AI to negotiate effectively, not just through random chance or simple rules, but through learned emotional strategy.
University of Cambridge · Technical University of Munich · Exiger LLC · The Alan Turing Institute
cs.CL, cs.AI
Submitted: 2026-05-26
Updated: 2026-09-03
Comments: Code: https://github.com/Yunbo-max/EmoDistill
Code: https://github.com/Yunbo-max/EmoDistill
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: The paper, "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation," addresses the challenge of equipping large language model agents with sophisticated
Key concepts
- Emotional Negotiation Skill
- This is defined as a reward-annotated turn that binds the dialogue state with its emotional stance. These skills are derived from an offline sweep where two large language models (LLM vs LLM) interact.
- EmoDistill Framework
- The overall system uses structured knowledge to teach AI agents strategic competence. It moves beyond simple, brittle prompt templates by allowing the agent to learn complex emotional dynamics necessary for effective negotiation.
- Decoupling Selection and Expression
- The authors separate the act of choosing an emotion from the act of expressing it. They train a smaller, targeted language model (SLM) specifically to execute that skill, rather than just instructing a large LLM to 'be angry.'
- Offline Skill Distillation
- This is the operational benefit of learning from existing data rather than running millions of live negotiations. It creates a stable foundation for building robust commercial AI tools by capturing complex strategic patterns.
Terminology
Summary
The paper, EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation,
addresses the challenge of equipping large language model agents with sophisticated behavioral skills required for high-stakes negotiation. It proposes a framework that explicitly models and controls emotional states during dialogue generation, arguing that controlling the affective layer is critical for achieving successful outcomes in adversarial settings.
The Joint Skill Pattern: Emotion-Conditioned Generation
The core mechanism trains the model to recognize and reproduce a specific joint pattern
within successful negotiation utterances. This pattern involves a labeled affective opener—such as “I’m terrified...” or “honestly, I’m a bit embarrassed...” —immediately followed by concrete, actionable information. The training process utilizes Supervised Fine-Tuning (SFT) cross-entropy on thousands of similar disaster turns to bias the LoRA adapter toward producing this two-piece structure at inference. This ensures that the generated response is not merely emotionally colored, but also structurally grounded in rationale and quantifiable details, such as a specific time frame or dependency (90 min, structural-scan dependency
).
Per-Turn Discrimination via Judge Policy Optimization (JPO)
The framework incorporates Judge Policy Optimization (JPO) to refine agent behavior by distinguishing between effective and ineffective turns within a single negotiation trajectory. JPO uses the per-turn judge advantage (rho t A t) as its policy-gradient update direction, allowing it to differentiate subtle tactical differences that standard sequence models miss. For instance, when comparing two turns leading to the same terminal reward, JPO learns to upweight a turn characterized by firm rejection + calibrated concession + leverage
while downweighting a turn exhibiting repetition + uncalibrated concession + no new leverage.
This per-turn discrimination is highlighted as the source of significant performance gains over SFT alone.
The Criticality of Decoupling IQL Selection from Generation
A key empirical finding tests the hypothesis that the emotion selection mechanism (IQL) is not a redundant layer. By comparing two configurations—(a) IQL+SFT+JPO and (b) Emotion-Free SFT+JPO—on held-out scenarios, the authors demonstrate that decoupling these stages severely degrades performance. The success rate on CRAD held-out scenarios drops significantly when the emotion block is stripped at inference. This reveals that while the LoRA adapter can generate both an angry anchor and a confused probe, in the absence of an explicit emotion call it has no signal about which mode to enter at any given state.
IQL as the Source of Negotiation Leverage
The Internal Quality Learner (IQL) selector is shown to be essential for unlocking specific leverage frames
embedded in the model's weights. Without the IQL, the adapter defaults to an emotion-marginalized mode,
which on CRAD is described as a conciliatory default that the counterparty does not feel pressure to move against.
Conversely, when the IQL guides state-conditional emotion calls (e.g., selecting ANGER first and then CONFUSION), it enables the agent to execute complex strategies—such as using ANGER to establish a hard anchor, followed by CONFUSION to force the opponent into justifying an overly extended timeline—thereby closing the deal where the emotion-free configuration fails due to exhausting its turn budget.
Improvements for AI systems
Based on this scientific paper extract, the core weakness of current LLM-based dialogue systems is their inability to reliably manage and leverage high-level strategic intent (emotion/rhetorical framing) alongside low-level linguistic generation.
The improvements must focus on decoupling strategic selection from tactical generation, ensuring that the model's output is not merely grammatically correct but strategically potent.
Here are the specific, actionable improvements for an AI system designed for high-stakes negotiation or complex dialogue simulation:
-
Improvement: The model architecture must integrate a dedicated, high-level Intent Quality Layer (IQL) that operates before the text generation phase. This layer selects the optimal rhetorical/affective stance (Emotion t) for the focal agent at each turn t.
-
Mechanism: Instead of allowing emotion to be implicitly learned during training, Emotion t must be explicitly selected from a defined set of strategic emotions (e.g., ANGER, CONFUSION, DISAPPROVAL, CURIOUSITY). This selection dictates the mode for the subsequent generation.
-
Technical Benefit: This prevents the model from defaulting to a
conciliatory default
(the emotion-marginalized mode) when high pressure is required, as demonstrated by the failure of Emotion-Free SFT+JPO. -
Improvement: The standard LoRA adapter training must be modified to accept and prioritize Emotion-Conditioned Prompts. The prompt structure must explicitly inject the selected Emotion t (e.g.,
[EMOTION: ANGER]) into the context window for the focal agent's turn. -
Mechanism: This forces the underlying language model to access and activate specific
leverage frames
embedded in its weights. The system learns thatANGER
requires a strong anchor, whileCONFUSION
requires a probing question demanding justification (as seen in debt 100). -
Technical Benefit: This maximizes the utility of the vocabulary learned during SFT+JPO, ensuring that the output is not just like an angry person speaking, but embodies the strategic function of anger (e.g., establishing an unyielding anchor).
-
Improvement: The fine-tuning loop must adopt a Per-Turn Judge Policy Optimization (JPO) update mechanism, rather than relying solely on terminal episode rewards.
-
Mechanism: The policy gradient update (rho t At) must weight the advantage (A t) of each individual turn based on its strategic impact (e.g., introducing new leverage, calibrating a concession).
-
Technical Benefit: This allows the system to differentiate between two equally successful trajectories: one that achieves the goal through high-leverage moves (Disapproval turn) and one that achieves it through repetition or uncalibrated concession (Annoyance turn). The model learns to avoid redundant or strategically weak turns, even if they lead to a positive outcome.
The resulting AI system moves beyond simply simulating conversation; it simulates strategic negotiation.
- Execute Multi-Phase Strategy Shifts:
- The system can execute complex, multi-turn plans that require deliberate shifts in rhetorical framing. For example, it can start by using ANGER to establish a high anchor (e.g., 24 days), then strategically pivot to CONFUSION later in the dialogue to force the opponent to justify their counter-proposal (152 days), thereby collapsing their position into accepting the initial anchor.
- Calibrate Concessions with Strategic Context:
- It avoids arbitrary concessions. When forced to concede, it calculates the concession not just based on proximity to a target, but based on the strategic value of that move relative to the current leverage state (e.g., only conceding if a prior concession by the counterparty was significant, as in CRAD Case Study 3).
- Maintain High-Stakes Consistency:
- In high-stakes domains (legal, medical, financial), the system can maintain domain-specific terminology and leverage points. It understands that
legal stage
orstructural scan
are not just words, but critical constraints that must be repeatedly referenced to heighten pressure and credibility throughout the dialogue.
- Systematic Failure Prevention:
- By coupling IQL selection with SFT+JPO, the system guarantees that even if a simple emotional state is selected (e.g., CONFUSION), the output will contain the necessary probing mechanism (
Could you clarify exactly why...
) rather than merely expressing confusion vaguely.
Abstract
Post-trained LLMs are often optimized to produce helpful, polite, and accommodating responses. In adversarial negotiation, however, such behavior can become a vulnerability: emotionally framed language may influence an agent's bargaining decisions in ways that conflict with its user's objectives. We therefore introduce EmoDistill, an offline framework for distilling emotional negotiation skills from LLM-LLM interactions into smaller language-model agents. Here, an emotional negotiation skill is a state-conditioned behavior that determines which explicit emotion to invoke in a bargaining state and how to realize that emotion as an effective negotiation utterance. EmoDistill learns these two components separately: an Implicit Q-Learning (IQL) selector learns which emotion to express in each bargaining state, while a LoRA-adapted 7B policy learns emotion-conditioned expression through Supervised Fine-Tuning (SFT) and Judge Policy Optimization (JPO). Across four emotion-sensitive negotiation domains, the full EmoDistill policy achieves competitive utility and improves over vanilla and IQL-only baselines in most settings. Emotion-free ablations show that removing the explicit emotion channel substantially reduces overall negotiation utility, while transfer experiments reveal partial, domain-dependent transfer and robustness to unseen LLM counterparties.
Sources
- Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
- Constitutional AI: Harmlessness from AI Feedback
- AgreeMate: Teaching LLMs to Haggle
- DeepSeek-V3 Technical Report
- Trustless Autonomy: Understanding Motivations, Benefits, and Governance Dilemmas in Self-Sovereign Decentralized AI Agents
- EmoDebt: Bayesian-Optimized Emotional Intelligence for Strategic Agent-to-Agent Debt Recovery
- EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation
- Qwen2.5 Technical Report
- Proximal Policy Optimization Algorithms
- ACE: A LLM-based Negotiation Coaching System
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Large Language Model Sentinel: LLM Agent for Adversarial Purification
- EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiation
- EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
- Qwen3 Technical Report
- DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
- Memento-Skills: Let Agents Design Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering