Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
cs.SD, cs.CV, eess.AS
Submitted: 2026-08-28
Updated: 2026-09-25
Code: https://github.com/project-numina/aimo-progress-prize
Project page: https://step-out.github.io/Motion-Omni-Page
Terminology
Sources
- Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- GLM-TTS Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- OpenThoughts: Data Recipes for Reasoning Models
- FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
- Kimi-Audio Technical Report
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
- MIBURI: Towards Expressive Interactive Gesture Synthesis
- GPT-4o System Card
- Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Qwen3-TTS Technical Report
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- Qwen2.5-Omni Technical Report
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment