EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
summary
The gist
The paper, "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation," addresses the challenge of equipping large language model agents with sophisticated
In short
The episode discusses EmoDistill, a method for training AI agents in adversarial negotiation using offline skill distillation. This approach teaches agents strategic emotional competence by distilling complex dynamics from high-performing examples. The system outperforms baseline LLMs and demonstrates cross-domain transferability, resulting in robust, strategically superior AI tools.
Key concepts
- Emotional Negotiation Skill
- This is defined as a reward-annotated turn that binds the dialogue state with its emotional stance. These skills are derived from an offline sweep where two large language models (LLM vs LLM) interact.
- EmoDistill Framework
- The overall system uses structured knowledge to teach AI agents strategic competence. It moves beyond simple, brittle prompt templates by allowing the agent to learn complex emotional dynamics necessary for effective negotiation.
- Decoupling Selection and Expression
- The authors separate the act of choosing an emotion from the act of expressing it. They train a smaller, targeted language model (SLM) specifically to execute that skill, rather than just instructing a large LLM to 'be angry.'
- Offline Skill Distillation
- This is the operational benefit of learning from existing data rather than running millions of live negotiations. It creates a stable foundation for building robust commercial AI tools by capturing complex strategic patterns.
Terminology used across episodes
This episode discusses
- EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation · Paper Radio
- Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
- Constitutional AI: Harmlessness from AI Feedback
- AgreeMate: Teaching LLMs to Haggle
- DeepSeek-V3 Technical Report
- Trustless Autonomy: Understanding Motivations, Benefits, and Governance Dilemmas in Self-Sovereign Decentralized AI Agents
- EmoDebt: Bayesian-Optimized Emotional Intelligence for Strategic Agent-to-Agent Debt Recovery
- EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation
- Qwen2.5 Technical Report
- Proximal Policy Optimization Algorithms
- ACE: A LLM-based Negotiation Coaching System
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Large Language Model Sentinel: LLM Agent for Adversarial Purification
- EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiation
- EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration
- Qwen3 Technical Report
- DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
- Memento-Skills: Let Agents Design Agents
The paper
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation · Read on arXiv
University of Cambridge · Technical University of Munich · Exiger LLC · The Alan Turing Institute
Post-trained LLMs are often optimized to produce helpful, polite, and accommodating responses. In adversarial negotiation, however, such behavior can become a vulnerability: emotionally framed language may influence an agent's bargaining decisions in ways that conflict with its user's objectives. We therefore introduce EmoDistill, an offline framework for distilling emotional negotiation skills from LLM-LLM interactions into smaller language-model agents. Here, an emotional negotiation skill is a state-conditioned behavior that determines which explicit emotion to invoke in a bargaining state and how to realize that emotion as an effective negotiation utterance. EmoDistill learns these two components separately: an Implicit Q-Learning (IQL) selector learns which emotion to express in each bargaining state, while a LoRA-adapted 7B policy learns emotion-conditioned expression through Supervised Fine-Tuning (SFT) and Judge Policy Optimization (JPO). Across four emotion-sensitive negotiation domains, the full EmoDistill policy achieves competitive utility and improves over vanilla and IQL-only baselines in most settings. Emotion-free ablations show that removing the explicit emotion channel substantially reduces overall negotiation utility, while transfer experiments reveal partial, domain-dependent transfer and robustness to unseen LLM counterparties.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation".
Jane: The paper was written by Yunbo Long, Haolang Zhao, Lukas Beckenbauer, Liming Xu and Alexandra Brintrup from University of Cambridge and Technical University of Munich and Exiger LLC and The Alan Turing Institute.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: So, how does EmoDistill actually work? It sounds like a complex pipeline to build these agents.
Jane: Essentially, the authors are using a clever distillation process to teach emotional skills without running millions of costly live negotiations, which is a huge operational win.
Lu: They define an "emotional negotiation skill" as something that binds the dialogue state and its emotional stance—it’s a reward-annotated turn from an offline LLM-vs-LLM sweep.
Meng: This entire process relies on three stages: first, using an Implicit Q-Learning selector to choose the emotion, then Supervised Fine-Tuning via LoRA to learn how the express it in a small model, and finally refining it with Judge Policy Optimization or JPO.
Jane: The authors are decoupling the *selection* of emotion from its *expression*, so they aren't just telling an LLM to "be angry"; they are training a smaller, targeted SLM to execute that specific skill.
Tom: It’s like creating a specialized agent that has been trained by watching and learning from top performers, which is incredibly powerful.
Lu: And Lalam points out that this system learns not just *what* to say but *when* to say it, ensuring the emotional framing is perfectly timed for maximum leverage.
Meng: The engineering benefit of this specific mechanism is that we are taking a large, complex skill and distilling it into a manageable student model rather than relying on brittle prompt templates.
Lalam: It allows us to build agents that possess a form of genuine strategic competence by using structured knowledge instead of hoping they just happen to be the right emotion.
Improvements: Tom: Moving from *how* it works to *what* it achieves, the results are genuinely impressive across all four domains.
Jane: The full EmoDistill policy consistently outperforms vanilla LLM and SLM baselines, which is not just a slight improvement but a substantial difference in success rate and utility.
Lu: What’s particularly interesting is that the gains aren't just from the high-level emotion selection; they are proving that optimizing *how* the expression—the utterance itself—is realized is crucial for achieving better outcomes.
Meng: Looking at Table one it clearly shows that combining those three components, IQL, SFT, and JPO, provides a synergistic effect that simply choosing the right emotion alone cannot achieve.
Tom: And it’s not just in one specific area; the fact that EmoDistillC trained on Credit Recovery performs well when applied to Student Sleep Scheduling demonstrates cross-domain transferability.
Jane: It does, and this confirms what the paper' suggests: emotional cues are not just a cosmetic style but a core component of the strategy itself.
Lu: We can also see some fascinating trade-offs in their findings; for example, IQL+SFT+JPO often takes more rounds than other methods, but that is because they are achieving significantly better value per turn.
Meng: The engineering takeaway here is that the system isn't just succeeding; it’s succeeding in a way that is strategically superior to simply rushing an agreement.
Lalam: This proves that we can build reliable, high-performing agents by learning complex emotional dynamics instead of relying on simple, brittle heuristics.
Conclusion: Tom: We've seen how EmoDistill works and what it achieves—it’s a powerful tool for teaching strategic negotiation in AI agents. But before we wrap up, let's hear the final thoughts from everyone.
Jane: I hope the authors’ findings are encouraging; if an AI can learn to be strategically effective through emotional framing, it has real potential for making complex negotiations much more efficient across industries.
Lu: I'm extremely excited about seeing these skills transferred to even more advanced applications, beyond just simple agent-to-agent scenarios; the scope of this work is massive.
Meng: My main hope is that this offline approach provides a stable foundation for building robust commercial AI tools that can handle complexity without the instability we see in current online reinforcement learning methods.
Lalam: I think the ultimate impact will be in how this changes our perception of what a truly competent autonomous agent is, moving beyond just efficiency to including strategic emotional intelligence.
Tom: These are all incredibly powerful visions; we appreciate you sharing your insights today on "EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation."
Final Goodbye: Jane: We've seen a framework that teaches AI to negotiate effectively, not just by chance, but through learned emotional strategy.
Lu: I agree; the fact they can use this offline training to unlock future possibilities is truly thrilling.
Meng: The ability EmoDistill provides for avoiding those costly live negotiations is what makes this approach scalable and practical.
Lalam: It allows us to build agents that feel and operate with genuine strategic depth, which will improve our relationship with AI assistants moving forward.
Tom: It's a framework that teaches AI to negotiate effectively, not just through random chance or simple rules, but through learned emotional strategy.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language