Evaluating the Capabilities of LLMs for Persuasive Dialogue
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: 19 pages, 7 figures, 4 tables
Code: https://github.com/RobinsonJI/Ev
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce Persuasio, a multi-agent dialogue platform grounded in a
Terminology
Abstract
Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce Persuasio, a multi-agent dialogue platform grounded in a formal argumentation-based theory of persuasion dialogues that adjudicates logical winners during free-text debates. Using this system, we generated 192 debates on a UK political topic between humans and LLMs, and evaluated 22 interlocutors through both automated adjudication and 9,702 crowdsourced pairwise judgements across 1, 386 annotation instances. We observed a consistent decoupling between subjective and formal persuasiveness: LLMs dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive. Multi-agent and retrieval-augmented variants further widened this divergence. These findings reveal a systematic gap between rhetorical fluency and formal argumentative strength in LLM-based persuasive dialogues.
Sources
- Susceptibility to Influence of Large Language Models
- Evidence of a log scaling law for political persuasion with large language models
- Large Language Models are as persuasive as humans, but how? About the cognitive effort and moral-emotional language of LLM arguments
- The Thin Line Between Comprehension and Persuasion in LLMs
- Training language models to follow instructions with human feedback
- Validating Political Position Predictions of Arguments
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering