MotionHiFlow: Text-to-motion via hierarchical flow matching

arXiv:2604.23264 · cs.CV · Submitted 2026-04-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MotionHiFlow: Text-to-motion via hierarchical flow matching".

Jane: MotionHiFlow proposes a hierarchical flow matching framework to generate 3D human motions progressively from low to high temporal scales,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we're diving into the paper "MotionHiFlow: Text-to-motion via hierarchical flow matching," and it looks like the core idea is using a hierarchical flow matching framework to generate three dee human motions that are both text-aligned and detailed <ref:2604.23264#pg0,MotionHiFlow: Text-to-motion via hierarchical flow matching>. Jane, can you give us the high-level summary of what they're proposing?

Jane: Absolutely, Tom. Essentially, MotionHiFlow tackles the issue where single-scale methods struggle with capturing both the overall meaning from text and the fine temporal details of a motion. The authors propose a hierarchical approach where they generate motions progressively, starting from low temporal scales to high ones. This lets the early stages focus on understanding the big picture semantics first before refining those finer details later on.

Lu: That hierarchical way of thinking really resonates with how complex systems operate, doesn't it? Thinking about motion as a hierarchy rather than one continuous stream makes sense for modeling human movement, Lu adds, and I'm excited to see how they structure that flow path from low to high temporal scales.

Meng: From an engineering standpoint, the idea of capturing high-level semantics early and then refining details sounds like it could make the training process much more stable than trying to learn everything at once. Does this hierarchical structure introduce any practical overhead in terms of computational cost during inference?

Lalam: I see potential here for improving how we understand human culture, Meng. If these models can accurately generate motions with rich semantics, it means we could create AI representations that capture nuanced human expressions far better than current methods.

Tom: That's a great point about stability, Meng. And the paper explains their mechanism through a stage-wise flow transformation S k that maps the start state of one stage to the end state of the next, which is a key part of this framework. Jane, can you elaborate on what those initial stages are supposed to capture specifically?

Jane: The abstract states that in the early stages, they aim to capture high-level semantics and coarse motion structures. This means they're focusing on getting the general shape and the overall action correct based on the text query before moving into finer temporal details. It’s about establishing that core textual alignment first.

Paper summary: Lu: The paper mentions that models trained solely on coarse motions often achieve robust semantic alignment, sometimes even outperforming those trained on fine-scale motions, which is a really interesting finding Lu adds, and it suggests that focusing too much on the fine details might actually hurt the model's ability to grasp the basic meaning of the text query.

Meng: That observation about coarse motions outperforming fine-scale ones is significant; it tells us where we should focus our initial training efforts if we want good results quickly. But what about how they handle that transition between these different scales, Tom?

Tom: That brings us to the novel cross-scale transition process, which is a big part of MotionHiFlow. They don't just upsample noisy data; they use a three-step process involving denoise, upsample, and renoise to link the start state of one stage with the end state of another across scales.

Jane: That cross-scale transition is where they ensure noise consistency is preserved across all those stages, which is crucial for maintaining coherence in the final output. It's not a simple interpolation; it’s a carefully formulated sequence designed to bridge the gap between scales without losing information.

Lalam: Preserving noise consistency sounds vital for generating realistic human motion sequences where every frame needs to flow smoothly into the next, Lalam adds, and that level of temporal coherence is something we need to improve for better cultural representation.

Tom: And they build this framework using several components, including a Motion VAE that uses Graph Convolutional Networks to capture topology, and a Text-Motion Diffusion Transformer that incorporates joint-aware positional encoding through Joint RoPE. Lu, what are your thoughts on how incorporating explicit human body topology into the motion representation helps with the flow matching?

Lu: The use of GCNs to capture human body topology allows the model to understand spatial dependencies among joints directly, which should inherently make the generated motions more physically plausible than just treating every joint independently. It gives the AI a spatial grammar to work with, and that's where I see immense creative potential for generating novel movement styles Lu adds.

Meng: On the practical side, if we can explicitly model the topology like that, it should translate into faster convergence during training because the AI isn't wasting cycles trying to learn how limbs connect in a general sense. But how do they manage those positional embeddings across different temporal and spatial contexts?

Paper summary: Jane: They address that with Joint RoPE, which adapts Rotary Position Embedding by integrating rotations derived from temporal displacement, relative spatial coordinates, and the kinematic tree structure. This seems like a very tailored way to embed skeletal motion information directly into the model's processing.

Tom: That kinematic tree structure sounds particularly clever for encoding the hierarchical relationships in human movement, Jane, and it ties back into that idea of understanding motion hierarchically we talked about earlier. Meng, does this level of structural detail translate into a more reliable output when generating complex actions?

Meng: It should do so by providing stronger constraints on the spatial arrangement of the joints during generation; if the model understands the kinematic tree, it's less likely to produce anatomically impossible poses that might trip up downstream applications, Meng adds.

Lalam: For culture and representation, if we can generate motions where the body structure is inherently respected at every temporal step, it means we are building AI that respects fundamental human form in its output Lalam notes.

Tom: So, to wrap up this overview of MotionHiFlow: Text-to-motion via hierarchical flow matching, Jane and I think the main point is that by using a hierarchical flow matching framework with a novel cross-scale transition process, they can generate motions that achieve both strong semantic alignment from text and rich fine-grained temporal details.

Jane: Exactly, Tom. It’s about moving beyond single temporal scales to capture the motion structure progressively, ensuring consistency across those stages so you get a more complete picture of the action.

Lu: The implications here are huge because it allows for the creation of highly nuanced synthetic human data that respects both semantic intent and physical constraints simultaneously, Lu adds. This opens up avenues for simulating complex human interactions with unprecedented fidelity.

Meng: From a deployment standpoint, if this framework can generate high-fidelity motion efficiently, it could be useful in creating realistic training data for robotics or virtual reality environments where movement accuracy is paramount.

Lalam: I think the real impact is how we use these detailed motions to build AI systems that are more capable of understanding and responding to human behavior in a more realistic way, Lalam concludes.

Paper summary: Tom: So, looking ahead at the title "MotionHiFlow: Text-to-motion via hierarchical flow matching," Jane and I think the implication is that for text-to-motion generation to become truly useful, we need methods that respect the inherent hierarchical nature of human movement structure.

Jane: That's a good summary, Tom. It moves away from treating motion as a flat sequence and embraces its natural structure, which is what makes it more powerful for capturing detailed actions.

Lu: I think the future work will likely explore how this framework can be adapted to even higher levels of abstraction or perhaps integrated with other modalities to generate richer scenes Lu adds.

Meng: As an engineer, I'm curious if they address the limitations of the current approach explicitly; specifically, what are they saying about where this method might struggle, Meng asks.

Tom: They do mention that while this approach is very effective on benchmarks like HumanMLthree dee and KIT-ML—where they achieved a highest R-Precision of zero point five six three and the lowest FID of zero point zero three two—the paper itself suggests that the complexity of motion generation still presents challenges, implying there's room for refinement in handling extremely rare or very intricate movements Lu adds.

Jane: That's fair, Tom. The authors acknowledge that while their method is strong on those metrics, it’s not perfect across every possible motion scenario yet.

Lalam: For us at the start-up side, the implication is that we can build AI systems that don't just generate motions but understand the underlying structure of human action better, Lalam adds.

Tom: So to wrap up this discussion on MotionHiFlow: Text-to-motion via hierarchical flow matching, we’ve seen how this hierarchical approach using cross-scale transitions allows for better semantic alignment and fine details than single-scale methods.

Jane: It really shows that breaking down a complex generation task into sequential, structured steps yields better results when the structure itself is meaningful.

Lu: The potential for simulating complex human interactions with high fidelity is where I see the most exciting creative possibilities lying Lu adds.

Meng: For practical application, the stability offered by this hierarchical flow matching suggests it could lead to more robust generative models in fields like motion capture or realistic digital avatar creation, Meng says.

Lalam: And ultimately, this kind of advance means our AI can create representations of human activity that are far more faithful and useful for understanding how people move through the world Lalam adds.

Conclusion: Tom: So, we've seen how MotionHiFlow uses hierarchical flow matching to tackle motion generation from text and get those fine details right. Jane, what's your take on that title and who came up with this work?

Jane: I think "MotionHiFlow" really captures the essence of what they did; it’s about a smooth, structured flow for human movement. The authors put together a sophisticated mix of diffusion models and VAEs to achieve this, which is pretty impressive.

Lu: It's fascinating that they combined flow matching with topology awareness from the Motion VAE; that shows a real deep dive into how we can model complex spatial relationships in movement data.

Meng: From an engineering standpoint, I'm curious how they managed to keep the noise consistent across all those stages without making the process computationally prohibitive during training.

Lalam: For me, the core idea is that by respecting the natural hierarchy of human motion structure, this AI can generate representations of culture and behavior that feel much more authentic.

Tom: Exactly! It's not just about generating a video; it's about generating meaningful motion guided by text in a way that respects the body's underlying physical rules.

Jane: And the authors managed to achieve very strong results on established benchmarks, which gives us solid proof that this hierarchical approach actually works for complex human actions.

Lu: I think what's most exciting is how they use those cross-scale transitions to bridge the gap between coarse semantic understanding and minute temporal details seamlessly.

Meng: That seamlessness is crucial for practical applications; if the transition were messy, we wouldn't get reliable output in a real-time system.

Lalam: It means our models can start understanding not just what a person is doing, but why they are doing it in a culturally relevant way, which has huge implications for how we build interactive digital experiences.

Tom: So, the conclusion is that MotionHiFlow proves you can structure motion generation hierarchically to get better semantic control and detail than older single-scale methods. Where do we go from here?

Sun Yat-sen University · Shandong University

cs.CV

Submitted: 2026-04-25

Updated: 2026-04-25

Comments: accepted to CVPR 2026

Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 9352-9363

Code: https://github.com/ai-lh/MotionHiFlow

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: MotionHiFlow proposes a hierarchical flow matching framework to generate 3D human motions progressively from low to high temporal scales, addressing the limitations of single-scale methods by

Key concepts

Hierarchical Flow Matching Framework
This framework generates motion in multiple stages, starting with coarse structure and text alignment in early stages, then refining temporal details. It uses a stage-wise flow transformation to map the start state of one stage to the end state of the next, ensuring progressive refinement.
Cross-Scale Transition Process
Instead of simple upsampling, this process links different motion scales by denoisng lower scales via extrapolation, upsampling higher scales with interpolation, and renoising noise at the higher scale. This ensures that noise consistency is maintained across all generated stages.
Motion VAE
This component uses Graph Convolutional Networks (GCNs) to capture the human body's topology explicitly. It encodes complex motion data into a compact latent representation by downsampling spatial graph information, which helps in learning robust representations of human movement.

Terminology

Summary

MotionHiFlow proposes a hierarchical flow matching framework to generate 3D human motions progressively from low to high temporal scales, addressing the limitations of single-scale methods by capturing high-level semantics first and then refining details. The core contribution is a novel cross-scale transition process that ensures noise consistency across stages, integrated with a Text-Motion Diffusion Transformer and a topology-aware Motion VAE to achieve superior semantic alignment and fine-grained motion details on benchmarks like HumanML3D and KIT-ML.

The gist

MotionHiFlow is a hierarchical flow matching framework that progressively generates motion from low to high temporal scales, achieving strong semantic alignment and rich fine-grained motion details.

Hierarchical Flow Matching Framework

The framework operates in the latent space encoded by a Motion VAE and progresses through multiple stages, where early stages focus on coarse motion structure and text alignment, while later stages refine temporal details. The process involves defining a stage-wise flow transformation Sk that maps the start state of one stage to the end state of the next. This is achieved by defining the start state as:

Start: x(k)tk−1 = (1 − tk−1)f(x0, rk) + tk−1f(f(x1, rk−1), rk/rk−1), where f(x, r) means performing a temporal resampling on x with a factor r. The end state is defined as:

End: x(k)tk = (1 − tk)f(x0, rk) + tkf(x1, rk).

Cross-Scale Transition Process

Instead of directly upsampling lower-scale noisy data, MotionHiFlow introduces a novel cross-scale transition process to link the start state of stage k+1 with the end state of stage k across scales. This process consists of three steps: 1) denoise: constructing the clean data at the lower scale by extrapolation; 2) upsample: generating clean data at higher scale with upsampling; and 3) renoise: constructing the noise data at the higher scale via interpolation. The transition is formally governed by equations that ensure noise consistency can be preserved across stages.

Model Architecture Components

The framework integrates several key components for efficient and stable generation:

  1. Motion VAE: This VAE utilizes Graph Convolutional Networks (GCNs) to explicitly capture human body topology, encoding input motion into a compact latent representation x ∈ Rl×j×d. It employs spatial graph downsampling (from J to j latent joints) using methods such as averaging or learnable pooling.

  2. Text-Motion Diffusion Transformer (TMDiT): This model harnesses hierarchical flow matching and explicitly incorporates structural dependencies among human joints through joint-aware positional encoding (Joint RoPE). TMDiT processes the noisy motion latent x(k)t and conditioning word-level text embedding c, fusing information about the current timestep t, sentence-level text embedding cvec, and staged scale rk into an embedding y that modulates the TMDiT blocks.

  3. Joint RoPE: This mechanism adapts Rotary Position Embedding for skeletal motion generation by integrating rotations derived from temporal displacement, relative spatial coordinates, and the kinematic tree structure. It enforces skeletal symmetry by applying 1D RoPE segments to encode temporal position, 2D spatial coordinates relative to the pelvis in a T-pose, and depth in the kinematic tree.

Training and Inference

The training pipeline is conducted in two stages. The first stage trains the Motion VAE by minimizing a composite objective that combines standard VAE losses (reconstruction and KL loss) with an auxiliary term that improves temporal robustness, specifically minimizing Laug = ∥Dec(f(x, r)) − f(M, r)∥2. The second stage freezes the Motion VAE and trains the TMDiT model vθ using the hierarchical flow-matching loss (Eq. 6). Classifier-free guidance (CFG) is applied during training by randomly replacing the text condition c with a null token ∅ at a 10% probability. Inference follows Algorithm 1, which involves sampling initial noise x0, initializing the start state xˆ(k)0 ← f(x0, r1), and iteratively performing the flow transformation and cross-scale transition steps for each stage k=1 to K. The final clean sample is obtained by decoding the resulting latent representation xˆ1 through the Motion VAE decoder.

Evaluation

The framework was evaluated on HumanML3D and KIT-ML datasets using R-Precision, Multimodal Distance, and Frechet Inception Distance (FID). Quantitative results demonstrated that MotionHiFlow achieved state-of-the-art performance, recording the highest R-Precision (0.563 / 0.482) on HumanML3D and the "lowest FID (0.032 / 0.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems by integrating the concepts from MotionHiFlow, and what those improved systems will be able to do:


)MotionHiFlow Framework Integration

  1. Improve semantic alignment and long-term temporal coherence in Text-to-Motion generation by adopting a hierarchical, coarse-to-fine generation strategy.

  2. Implement a multi-scale generative process that explicitly models high-level semantic structures first (low temporal scale) before refining fine details (high temporal scale).

  3. Introduce a novel cross-scale transition process (denoise, upsample, renoise) to link flows across scales, ensuring noise consistency and preventing the degradation of generation quality often seen in naive upsampling methods.

  4. Incorporate explicit structural dependencies among joints into the motion generation model using a Text-Motion Diffusion Transformer (TMDiT) combined with Joint RoPE positional encoding.

)Improved AI System Capabilities

The resulting AI system will be able to perform the following:

  1. Generate highly realistic 3D human motions that are precisely aligned with complex, nuanced natural language descriptions (e.g., A person walks a circle or A person runs straight ahead, then makes a 90 degree left turn and walks slowly for several steps).

  2. Achieve superior temporal coherence over long sequences by ensuring the motion maintains high-level structural integrity across different temporal scales.

  3. Produce fine-grained, physically plausible motion details (e.g., specific limb movements, subtle rotations) that are not achievable by models operating at a single temporal resolution (addressing the limitation of current state-of-the-art).

  4. Demonstrate robust performance across diverse human action benchmarks like HumanML3D and KIT-ML, achieving state-of-the-art metrics in semantic alignment (R-Precision), motion realism (FID), and multimodal distance.

Sources

Related papers