SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling
cs.CV, cs.CL, cs.GR
Submitted: 2026-08-25
Updated: 2026-08-28
Code: https://github.com/OMEGA-i/SeMoCo-Tokenizer
Terminology
Sources
- AudioLM: a Language Modeling Approach to Audio Generation
- Executing your Commands via Motion Diffusion in Latent Space
- Simple and Controllable Music Generation
- High Fidelity Neural Audio Compression
- Qwen3-TTS Technical Report
- Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
- MotionGPT: Human Motion as a Foreign Language
- VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
- Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
- Kimodo: Scaling Controllable Human Motion Generation
- SOMA: Unifying Parametric Human Body Models
- Human Motion Diffusion Model
- Neural Discrete Representation Learning
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- MotionBricks: Scalable Real-Time Motions with Modular Latent Generative Model and Smart Primitives
- UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
- HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
- SoundStream: An End-to-End Neural Audio Codec
- SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models