CogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation
cs.MM, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 17 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Persuasive video generation (PVG) is a valuable yet under-explored research topic.
Terminology
Abstract
Persuasive video generation (PVG) is a valuable yet under-explored research topic. Despite the significant advances in multimodal content generation, AI-empowered automated creation of human-made-like videos with substantial persuasiveness remains a formidable challenge. In this paper, we propose CogenPVG, a novel Cognitive-Enhanced reflective multi-agent framework tailored for Persuasive Video Generation task. Given the topic and stance from the user, we decouple the sophisticated generation process into four sequential stages: argument reasoning, storyboard planning, asset creation, and post-editing, imitating the workflow of human video producers. To ensure high persuasiveness, each stage is equipped with a pair of generator and critic agents, following a reflective refinement scheme grounded in a solid psychological theory of persuasion, the Elaboration Likelihood Model (ELM). In the argument reasoning stage, we generate highly logical and credible reasoning thoughts under the guidance of critical thinking theory, enabling cognitive enhancement via the central route of the ELM. For the other three stages, we generate and optimize multimodal assets, assembling them into a persuasive video guided by theories of heuristics, as the peripheral route of the ELM. To the best of our knowledge, CogenPVG is the first work focused on general persuasive topics, without being confined to commercial purposes. Extensive experiments and comprehensive analysis demonstrate that our framework achieves the best persuasion performance, thereby proving the effectiveness of our proposed multi-agent framework for the PVG task.
Sources
- GPT-4 Technical Report
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
- PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings
- DeepSeek-V3 Technical Report
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
- Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
- Gemini: A Family of Highly Capable Multimodal Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
- Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
- Automated Movie Generation via Multi-Agent CoT Planning
- Qwen3-Omni Technical Report
- MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression