BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://wengwanjiang.github.io/BiMoGen-Page
Terminology
Sources
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Continuous diffusion for categorical data
- Scaling Diffusion Language Models via Adaptation from Autoregressive Models
- Classifier-Free Diffusion Guidance
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- ReMoMask: Retrieval-Augmented Masked Motion Generation
- VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
- Decoupled Weight Decay Regularization
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
- Absolute Coordinates Make Motion Generation Easy
- Large Language Diffusion Models
- UniMo: Unified Motion Generation and Understanding with Chain of Thought
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
- UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
- FineMotion: A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing
- MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
- DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models