EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation
cs.GR, cs.CV, cs.LG, cs.MM, cs.RO, cs.SD, eess.AS
Submitted: 2026-08-19
Updated: 2026-08-19
Journal ref: The 1st International Workshop on Joint Audio-Video Comprehension and Generation (JAV-CG), co-located with ACM Multimedia 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state.
Terminology
Abstract
Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produce only linguistic outputs, leaving a critical gap in embodied response generation. We identify and address a failure of emotion conditioning: like other conditional generators that under-use weak conditioning signals, a flow-matching model given both a rich audio embedding and a discrete emotion label suppresses the emotion, generating near-identical motion regardless of the specified emotion. We present EMODY Flow, a lightweight (around 35M parameters) flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions. A training-time auxiliary emotion classifier restores emotion sensitivity by forcing generated motion to be emotion-identifiable. EMODY Flow sets a new state of the art on BEAT2 gesture quality, with FGD 0.302, Beat Correlation 0.853, and Diversity 24.62 - improving over the best prior results by 26%, 5%, and 62% respectively - and transfers to zero-shot facial animation on TFHP without domain-specific fine-tuning. Beyond these quantitative gains, the classifier yields clearly emotion-separated motion, which we demonstrate qualitatively through a multidimensional-scaling analysis of the generated gestures.
Related papers
- SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
- CADReasoner: Iterative Program Editing for CAD Reverse Engineering
- QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning
- DrawVideo: Grounded and Faithful Multi-Shot Video Generation from Storyboard Keyframe Sketches
- MotionPersona: Real-Time Locomotion Control across Personas, Bodies, and Styles
- MeshSplatBench: A Unified Benchmark for Triangle- and Mesh-Based Neural Rendering