Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
cs.CV
Submitted: 2025-12-30
Updated: 2026-09-16
Project page: https://chenhaoqcdyq.github.io/LMR
Terminology
Sources
- Human Motion Diffusion Model
- MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space
- MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
- RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
- Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
- Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Training Large Language Models to Reason in a Continuous Latent Space
- (Social) Trouble on the Road: Understanding and Addressing Social Discomfort in Shared Car Trips
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- Visual Chain-of-Thought Diffusion Models
- Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
- Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
- LLMs as Layout Designers: Enhanced Spatial Reasoning for Content-Aware Layout Generation
- Thinking with Generated Images
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
- PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation
- Fast Transformer Decoding: One Write-Head is All You Need
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models