EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

arXiv:2609.16011 · cs.GR, cs.CV, cs.LG, cs.MM, cs.RO, cs.SD, eess.AS · Submitted 2026-08-19 · Read on arXiv

cs.GR, cs.CV, cs.LG, cs.MM, cs.RO, cs.SD, eess.AS

Submitted: 2026-08-19

Updated: 2026-08-19

Journal ref: The 1st International Workshop on Joint Audio-Video Comprehension and Generation (JAV-CG), co-located with ACM Multimedia 2026

DOI: 10.1145/3840475.3841439

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state.

Terminology

Abstract

Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produce only linguistic outputs, leaving a critical gap in embodied response generation. We identify and address a failure of emotion conditioning: like other conditional generators that under-use weak conditioning signals, a flow-matching model given both a rich audio embedding and a discrete emotion label suppresses the emotion, generating near-identical motion regardless of the specified emotion. We present EMODY Flow, a lightweight (around 35M parameters) flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions. A training-time auxiliary emotion classifier restores emotion sensitivity by forcing generated motion to be emotion-identifiable. EMODY Flow sets a new state of the art on BEAT2 gesture quality, with FGD 0.302, Beat Correlation 0.853, and Diversity 24.62 - improving over the best prior results by 26%, 5%, and 62% respectively - and transfers to zero-shot facial animation on TFHP without domain-specific fine-tuning. Beyond these quantitative gains, the classifier yields clearly emotion-separated motion, which we demonstrate qualitatively through a multidimensional-scaling analysis of the generated gestures.

Related papers