GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression
cs.CV, cs.GR, cs.HC
Submitted: 2026-09-18
Updated: 2026-09-18
Project page: https://andypinxinliu.github.io/GestureFAR
Terminology
Sources
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
- DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
- One Step Diffusion via Shortcut Models
- Autoregressive Image Generation without Vector Quantization
- Flow Matching for Generative Modeling
- TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
- BEAT: A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis
- SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
- Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Towards Variable and Coordinated Holistic Co-Speech Motion Generation
- Consistency Models
- MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models