MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation
cs.CL, cs.AI, cs.CY, cs.HC, cs.MA
Submitted: 2025-09-30
Updated: 2026-09-09
Comments: Accepted to EMNLP 2025 Industry Track (https://aclanthology.org/2025.emnlp-industry.26.pdf)
Code: https://github.com/openai/evals
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Evaluating Task-oriented Dialogue Systems: A Systematic Review of Measures, Constructs and their Operationalisations
- Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms
- On the Worst Prompt Performance of Large Language Models
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
- COOPER: Coordinating Specialized Agents towards a Complex Dialogue Goal
- Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues
- MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents
- LLMs Get Lost In Multi-Turn Conversation
- IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
- LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Self-Refinement of Language Models from External Proxy Metrics Feedback
- Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset
- Contrastive Speaker-Aware Learning for Multi-party Dialogue Generation with LLMs
- Muse: A Multimodal Conversational Recommendation Dataset with Scenario-Grounded User Profiles
- PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
- SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
- Exploring the Impact of Personality Traits on Conversational Recommender Systems: A Simulation with Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering