ECHO-G: Embodied Co-speech Humanoid mOtion Generation

arXiv:2609.39575 · cs.RO, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "ECHO-G: Embodied Co-speech Humanoid mOtion Generation".

Dev: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, to recap what we've seen so far, ECHO-G is essentially proposing a system that generates full-body robot motion by jointly conditioning on the speech audio and the timed transcript. The core claim of this paper is that this joint conditioning allows the SpeechGrounded Diffusion Transformer, or SGDiT, to effectively combine frame-aligned acoustic features with linguistic context while maintaining their individual strengths.

Dev: Right, Rosa; it’s claiming that this approach models the one-to-many relationship between an utterance and the corresponding motion directly in robot space. They argue that existing methods often only look at acoustics or text separately, but ECHO-G integrates them into a unified generative model for motion generation.

Taro: What matters here is the integration itself; they’re not just stacking two models on top of each other; they are fusing the acoustic and linguistic conditions within each transformer block before value aggregation, which suggests a deeper level of interaction between those different modalities during the generation process.

Rosa: Precisely, Taro; that fusion is key because it preserves the distinct granularities of both inputs—the fine details in how a robot sounds versus the specific meaning conveyed by individual words. It’s about making sure the prosody and the content are perfectly aligned physically.

Dev: From my perspective as an engineer, this unified model structure is promising because it should provide a more coherent output than separate systems that might just try to stitch together motion derived from audio and motion derived from text independently. I’m looking for that coherence in the generated sequence.

Taro: Coherence in physical execution is where I live; if the system generates motions that feel physically awkward or mismatched because the linguistic context doesn't align with the acoustic timing, then it fails its autonomy purpose regardless of how good its internal math looks.

Rosa: That’s a fair point; it moves beyond just generating plausible-looking motion to generating motion that is actually communicative in a human-robot sense. It’s about making the robot look and move like it’s truly speaking the right thing at the right time, which is what this ECHO-G framework aims to achieve.

Dev: So, if I summarize the thesis simply, ECHO-G proposes using SGDiT trained with rectified flow matching to directly model that complex one-to-many relationship between speech and full-body motion references within a robot's coordinate system. It’s about predicting exactly what the robot should do given what it’s hearing and reading.

Taro: And the fact that they released a BEAT2-derived dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency gives us concrete things to actually test against this framework when we look at its real-world potential.

Rosa: Exactly; that data package is what moves this from a theoretical concept to something we can evaluate practically, showing us if the motion quality metrics actually correlate with how natural the co-speech interaction feels to a human observer.

Dev: So, it boils down to having a model that isn't just making motions based on audio alone or text alone, but one that understands the interplay between them in generating motion references for a humanoid robot. That’s what this paper is about.

Taro: And I think the implications are significant because if we can get this level of coordination right, we start bridging the gap between simple commands and truly embodied conversation for robots.

Rosa: Indeed, Taro; it suggests a future where robots don't just follow explicit commands but can participate in more nuanced, natural-sounding exchanges with people. That’s what's exciting about the direction this research is heading.

Dev: And from an engineering standpoint, if the runtime is competitive with other methods while maintaining those quality scores on that dataset, then it becomes a viable candidate for practical integration into existing humanoid platforms.

Conclusion: Rosa: So, wrapping up the discussion on ECHO-G, we’ve looked at how this framework uses SGDiT to generate motion references conditioned on speech audio and text transcripts. The authors are Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, and Hao Xu.

Dev: Yes, the implications seem to point toward a system that can create much more natural interactions for humanoids by aligning physical movement with spoken language in a way that goes beyond simple pre-programmed actions. It’s about achieving a level of embodied co-speech generation that feels genuinely communicative.

Taro: I think the real impact is in how it helps us understand the underlying dynamics of what makes human-robot communication feel natural; it gives us a blueprint for how to structure AI systems to exhibit more nuanced physical responses when the world throws unexpected inputs at them.

Rosa: It gives us a concrete way to measure that naturalness, which is something we’ve struggled with before; we now have metrics like FGD and Div that tell us if the motion feels right in terms of rhythm and content alignment. It helps us move toward building robots that can engage in more complex physical dialogues.

Dev: From an engineering standpoint, the fact that they established a methodology for training these models using rectified flow matching provides a solid mathematical foundation for building reliable, controllable generation pipelines. That’s a big step for engineers trying to deploy things reliably.

Taro: And looking forward, I see this work setting the stage for exploring more expressive behavior where we can finally teach robots gesture-aware movements that are truly contextually relevant, not just random movements.

Rosa: That's the long-term vision; moving toward a future where robots can have richer, more meaningful physical interactions based on what they hear and read. It’s about giving them a body that reflects their intelligence in real time.

Yizhao Li, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang

Nanjing University · Beihang University · The Hong Kong University of Science and Technology (HKUST) · Mondo Robotics

cs.RO, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 8 pages, 5 figures, 3 tables. Project page: https://echo-g-project.github.io/

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion.

Key concepts

SpeechGrounded Diffusion Transformer (SGDiT)
This is a core component that maps noisy motion sequences to desired movements. It combines acoustic features from speech audio with token-level linguistic context from text transcripts. This allows the model to understand how specific sounds and words should translate into physical body motions in robot space.
Transcript Cross-Attention Mechanism
This mechanism retrieves linguistic context from the transcript using two paths: a global path for overall content access and a local path that prioritizes temporally nearby tokens. This ensures the generated motion respects both the complete meaning of the speech and its specific timing.
Rectified Flow Matching
This is the training objective used to teach SGDiT. Instead of standard training, it matches a target flow velocity derived from desired motion with actual frame differences. This method effectively trains the model to generate smooth, realistic motion sequences that align precisely with the speech.
Robotspace Dataset (BEAT2-derived)
The authors created a dataset by retargeting existing data and filtering it for embodiment quality. This dataset is used to train and test ECHO-G, focusing on co-speech characteristics, robot motion quality, and runtime efficiency for practical deployment.

Terminology

Summary

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. The gist: ECHO-G presents a framework jointly conditioned on speech audio and timed transcripts to generate full-body robot motion references using a SpeechGrounded Diffusion Transformer (SGDiT).

ECHO-G Framework Overview

ECHO-G is designed to equip speaking humanoids with body motion that reflects both how an utterance is spoken and what it conveys, while respecting the robot’s embodiment. The framework jointly conditions on speech audio and timed transcripts. Its core component is the Speech-Grounded Diffusion Transformer (SGDiT), which combines frame-aligned acoustic features with token-level linguistic context to preserve their distinct granularities. This approach models the one-to-many utterance–motion relationships directly in robot space.

SGDiT Architecture and Conditioning

The SGDiT maps a noisy normalized motion sequence and conditions it using the combined audio and text input. The generator operates on normalized motion features, where each frame's representation incorporates a projected motion frame, its aligned acoustic condition, and a positional embedding: zτ = Wrxt,τ +Aτ +pτ. Temporal conditioning is achieved by combining acoustic features (mapped to A) and linguistic context (H), derived from token embeddings of the transcript.

Transcript Cross-Attention Mechanism

The model utilizes a cross-attention mechanism to retrieve linguistic context through two paths over the same transcript features. The global path provides content-based access to the full token sequence, while the local path adds a preference for temporally nearby tokens, yielding temporally reweighted attention: gτ = α max n∈I exp(bτn), Πτn = (1−gτ)Πg τn +gτΠl τn. This mechanism ensures that the motion generation is guided by both the overall content and the temporal structure of the utterance.

Training Objective and Flow Matching

SGDiT is trained with rectified flow matching to model the one-to-many relationship between utterances and full-body robot motion. The training objective involves matching a target flow velocity, where Vˆ = vθ (Xt,t,c) is supervised by matching the target flow and its adjacent-frame differences: Lflow = E[MSE(Vˆ,V⋆)], Ltemp = E[MSE(∆τVˆ,∆τV⋆)], Lgen = Lflow +λtempLtemp.

Dataset and Evaluation

To support training and evaluation, the authors introduce a BEAT2-derived robotspace dataset through retargeting and embodiment-specific quality filtering. The benchmark covers co-speech characteristics, robot-motion quality, and runtime efficiency. In quantitative results, joint audio–text conditioning achieves the lowest scores across all co-speech metrics (FGD, ∆Div, ∆BA) and demonstrates superior motion quality compared to unimodal variants. Furthermore, a user study showed that joint audio–text conditioning received the highest overall mean rating for humanlikeness and rhythm matching.

Real-Robot Deployment

The framework is demonstrated on a physical humanoid robot using a fixed SONIC motion tracker to execute the generated joint-position references while playing the corresponding speech audio. This deployment confirms the capability of direct prediction in robot space for real-world execution. The system achieves an end-to-end inference time that is competitive with other methods, supporting its practical viability.

Limitations and Future Work

Current limitations include a lack of consistent translation between gains in co-speech characteristics and improvements in robot motion quality metrics like Jerk gap or foot-contact measures. Additionally, the model learns broad associations without explicit deictic gesture training, which may constrain instruction-aware generation. Future work plans include jointly improving gesture expressiveness and robot-motion consistency, enriching data for instruction-aware gestures, and extending the framework to causal streaming generation.

Conclusion

ECHO-G successfully presents a full-body humanoid co-speech generation framework that utilizes SGDiT to generate robot motion references from speech audio and timed transcripts. The release of the dataset, benchmark, and code supports reproducible research in this domain. Future efforts will focus on enhancing expressive behavior and consistency in physical execution.


The gist

ECHO-G presents a framework jointly conditioned on speech audio and timed transcripts to generate full-body robot motion references using a SpeechGrounded Diffusion Transformer (SGDiT).

How it works

  1. The generator operates on normalized motion features, where each frame's representation incorporates a projected motion frame, its aligned acoustic condition, and a positional embedding: zτ = Wrxt,τ +Aτ +pτ.

  2. Temporal conditioning is achieved by combining acoustic features (mapped to A) and linguistic context (H), derived from token embeddings of the transcript.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems based on the ECHO-G framework, and what those improved systems will be capable of:


The core improvement lies in developing a generative model for full-body humanoid motion that is explicitly conditioned on both the acoustic prosody (speech audio) and the linguistic content (timed transcript), trained directly in robot space.

Here are the specific improvements:

  1. A new generation framework, specifically the Speech-Grounded Diffusion Transformer (SGDiT), which integrates frame-aligned acoustic features with token-level linguistic context via a rectified flow matching mechanism.

  2. The training of this SGDiT model using rectified flow matching to directly model the one-to-many relationship between utterances and full-body robot motion, enabling the sampling of diverse motions for a single utterance.

  3. The creation and release of a comprehensive dataset (BEAT2) pairing speech audio, timed transcripts, and corresponding full-body robot motion references, along with robust training, inference, and evaluation code.

The improved AI system (ECHO-G) can perform the following specific tasks:

  1. Generate diverse full-body humanoid motions that are perfectly synchronized with the prosody (rhythm) of a given spoken sentence while accurately reflecting the semantic content of the words being said.

  2. Produce motion references directly in robot space, which are then executed by a fixed whole-body motion tracker on a physical humanoid robot, allowing for real-time co-speech interaction.

  3. Achieve superior co-speech characteristics compared to unimodal methods (audio-only or text-only), resulting in motions that exhibit broader arm extensions and more fluid transitions than existing systems.

  4. Demonstrate high fidelity and low error in motion quality metrics (e.g., reduced body jerk, lower foot-ground error, and decreased contact sliding speed) when compared to human motion benchmarks.

  5. Provide a robust system for robotic instruction following, as the model is trained on speech-text pairings, potentially enabling the generation of specific deictic or instructional gestures when an utterance calls for them (addressing a limitation noted in Section V).

Abstract

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.

Sources

Related papers