GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots".
Dev: Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Now, let’s talk a bit more about the title and who came up with this work, "GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots." It really captures the essence of what they’re trying to do.
Dev: I think the title clearly signals that we're dealing with co-speech gestures, which is a common area for AI research, but they’ve added the crucial element of workspace conditioning.
Taro: From an autonomy researcher's view, that conditioning aspect suggests a move toward more physically aware interaction policies rather than just purely semantic ones.
Rosa: Exactly; it moves beyond just what the robot says to considering the physical space available for performing that movement, which is something we need when deploying robots outside of a clean lab setting.
Dev: And looking at the authors, Bosong Ding and his team seem to have put together a framework that bridges several complex areas: motion representation, diffusion models, and robotic retargeting.
Taro: Their background in different domains is what’s impressive because it suggests they could handle the complexity of unifying those six co-speech corpora effectively.
Rosa: They managed to combine datasets that naturally have incompatible representations into a single geometric representation by mapping them onto a common upper-body and hand skeleton.
Dev: That transformation into that one hundred twenty-three-dimensional motion representation is where the real technical heavy lifting happens, as it standardizes the input for their diffusion model.
Taro: I’m curious about how they handled those differences in joint trajectories when mapping everything onto that common skeleton; did those differences cause major issues?
Rosa: The authors claim that the mean differences in joint trajectories across all source-robot combinations stayed below one degree, which shows a very good level of consistency.
Dev: That level of consistency is what makes it viable for retargeting to different robot embodiments later on, ensuring the generated motion translates reliably.
Taro: If that translation is reliable, then we’re talking about systems that could actually work in complex real-world scenarios where the robot's physical structure varies widely.
Rosa: That reliability, combined with the workspace conditioning mechanism, is what makes GestAdapt a powerful tool for making robots adapt their behavior on the fly to physical constraints.
The paper's summary: Dev: Moving on to summarizing the core of "GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots," it boils down to using a workspace-conditioned diffusion Transformer model.
Rosa: That’s right; the system generates motion conditioned on speech, motion history, and a specified workspace box, which is the main mechanism for achieving that spatial awareness.
Taro: So, the AI isn't just generating a generic gesture anymore; it’s actively considering where it can physically move while simultaneously generating what it says.
Dev: Precisely; they embed the six workspace boundaries into the Transformer blocks through residual cross-attention, allowing each temporal position to pay attention to those boundaries throughout the denoising process.
Taro: That means if an obstacle appears, the model can adjust its intended movement based on those spatial cues embedded in its internal state during motion synthesis.
Rosa: It’s a sophisticated way of embedding physical context into the generation process rather than just applying a filter at the end of the generation pipeline.
Dev: They also used an analytic guidance term during sampling time to actively reduce residual wrist boundary violations, which is a key part of keeping the generated motion physically sound.
Taro: That analytic guidance term sounds like a safety net that helps prevent those small, unphysical movements from occurring during the actual motion synthesis.
Rosa: And this whole setup means they combine learned conditioning and active guidance to produce gestures that are both spatially aware and physically plausible.
The paper's improvements: Dev: Let’s look at the specific technical improvements proposed by GestAdapt to see exactly how they went from previous methods.
Rosa: The authors detail several key enhancements, including the unification of multiple datasets into a shared geometric representation and the use of that one hundred twenty-three-dimensional motion representation.
Taro: That unified representation is what allows them to handle the incompatibility between different source datasets by mapping them into a common space first.
Dev: Then they have their workspace conditioning mechanism, which involves learning six boundary embeddings combined with a learned identity embedding that feeds into the Transformer blocks for spatial conditioning during denoising.
Rosa: Plus, they introduced analytic guidance at sampling time to handle those residual violations without needing extra scene annotations or complex scene annotations for every scenario.
Taro: I think the analytic guidance term is particularly interesting because it addresses the issue of residual wrist violations directly during the reverse step equation calculation.
Dev: And finally, they use GMR, or Generalized Motion Representation, for robot retargeting to map that abstract motion into specific robot joint space using kinematic models and Jacobians.
Rosa: That final step ensures that even after all the sophisticated generation happens, the resulting motion is translated accurately into a robot's actual physical joint configurations while respecting its limits.
Conclusion: Rosa: So, to wrap up on this paper, we see that GestAdapt successfully integrates workspace awareness into co-speech gesture generation through conditioning on a prescribed wrist workspace.
Dev: It achieves this by using a diffusion Transformer model and analytic guidance during sampling to produce motions that respect those spatial boundaries while maintaining smoothness.
Taro: From an autonomy standpoint, this means we can expect robots to perform more context-aware actions where the gesture adapts dynamically to their physical surroundings without needing explicit pre-programming for every possible scenario.
Rosa: The implications are that these systems could be deployed in ways that significantly increase the perceived quality of human-robot interaction in real settings.
Dev: This paper suggests that treating target space as part of the generation process is superior to just correcting trajectories afterward, which is a way to improve the actual synthesized output.
Taro: It confirms that incorporating spatial constraints into the learned generator substantially alters the gesture to satisfy new spatial constraints while preserving perceived quality.
Rosa: Overall, "GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots" presents a solid framework for improving how we design these interactive systems.
Bosong Ding, *, , ,
Department of Intelligent Systems, Tilburg University · TU-Dresden, Dresden, Germany · School of Artificial Intelligence, China University of Mining and Technology-Beijing (CUMTB), Beijing, China
cs.RO, cs.AI, cs.HC, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
The gist: Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion.
Key concepts
- Unified Data Representation
- The framework combines six different gesture datasets (like BEAT and TED-Expressive) into a single, shared 123-dimensional geometric representation. This process aligns complex original recordings to a common upper-body and hand skeleton, ensuring the motion data is compatible for robot use without losing anatomical detail.
- Workspace Conditioning Mechanism
- The core is a diffusion Transformer model that generates motion conditioned on speech, history, and specific workspace boundaries (W). These boundaries are embedded as tokens that guide the generation process at every temporal step. This mechanism ensures the generated gestures are inherently aware of and adapted to the physical space they must occupy.
- Analytic Workspace Guidance
- During motion sampling, a mathematical guidance term is applied to penalize wrist positions that violate the prescribed workspace boundaries. This cost function guides the generation process toward valid space, significantly reducing unwanted residual violations in the final robot motion.
Terminology
Summary
Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. The GestAdapt framework introduces a workspace-conditioned approach that conditions co-speech gesture generation on a prescribed wrist workspace, allowing generated gestures to change with available space while preserving plausibility and smoothness.
The gist
GestAdapt is a workspace-conditioned framework that conditions co-speech gesture generation on a prescribed wrist workspace, learning from six complementary co-speech corpora through a shared motion representation and supporting retargeting to different robot embodiments.
Unified Data Representation
The framework begins by unifying multiple datasets into a shared geometric representation. The authors combine six datasets—BEAT, SHOW, Streamer, TED-Expressive, Trinity, and ZEGGS—which differ in recording settings and gesture distributions. To handle incompatible native representations, the motion is first evaluated on each original kinematic hierarchy before extracting anatomically corresponding Cartesian landmarks that are then aligned to a common upper-body and hand skeleton. This results in a 123-dimensional motion representation at 15 Hz, which preserves the geometry required for robot retargeting, with mean differences in joint trajectories remaining below 1° across source-robot combinations.
Workspace Conditioning Mechanism
The core of GestAdapt is a workspace-conditioned diffusion Transformer model. This model generates future motion conditioned on speech, motion history, and a specified workspace box (W). The six boundaries of the workspace are embedded by a shared MLP and combined with a learned boundary identity embedding. After temporal self-attention, each Transformer block injects these workspace tokens through residual cross-attention:
Z(l) = Ze(l) + CrossAttnl LN(Ze(l), C, C), where Ze(l) contains the temporal motion features and C ∈ R(6×D) contains the six workspace tokens. This allows each temporal position to attend to the workspace boundaries throughout denoising.
Analytic Workspace Guidance
To reduce residual wrist violations that can occur during sampling, GestAdapt employs an analytic guidance term at sampling time, following the idea of testtime guidance in diffusion-based motion generation. For each future frame t, wrist s, and spatial axis a ∈ (x,y,z), the workspace violation is defined as:
v t,s,a = max(la − wbt s,a, wbt s,-a − ua t).
The guidance cost is calculated as:
CW = (1/6T) ∑ t,s,a v squared + max t,s,a v squared.
The sampling process uses the reverse step equation: eεk = εk + λk∇XkCW. The guidance term is set to zero when G=0, which retains workspace conditioning without analytic guidance but reduces residual violations.
Robot Retargeting and Execution
The diffusion model generates motion in a unified body representation, and then the motion is transferred to the target embodiment using GMR (Generalized Motion Representation). GMR maps generated upper-arm and forearm directions and palm orientations into the robot frame using embodiment-specific coordinate calibration. Finally, GMR solves for robot joint configurations using the robot kinematic model, Jacobians, and joint limits. The requested workspace is registered in the robot base frame and converted to a metric box using an embodiment-specific scale. The execution involves solving a constrained optimization problem:
min ∆q s ≥ 0
1/2 ∆q T HIK ∆q + c T s
s.t. AW ∆q - s ≤ bW, Ahard∆q ≤ bhard,
where the diagonal weight matrix Λ penalizes workspace slack to soften the constraints. Workspace constraints are soft, while joint ranges and self-collision constraints remain hard physical constraints.
Evaluation and Results
Quantitative evaluation shows that generated motions remain close to the real-motion distribution while respecting the workspace; for instance, in a user study under modified workspace constraints, gestures received a mean quality score of 3.24/5, which was above the no-workspace variant (2.43/5) and below reference motions (3.68/5). In a real robot evaluation, motions generated with GestAdapt ranked first in 69.7% of comparisons when compared to the no-workspace variant and retargeted ground-truth motions constrained afterward, supporting the idea that treating target space as part of gesture generation is superior to modifying trajectories afterward. Ablation studies confirm that multi-corpus training improves workspace compliance, the analytic guidance further reduces residual violations, and a factorized motion representation improves both performance metrics. The user study further validates that incorporating the target workspace directly into the learned generator substantially alters the gesture to satisfy new spatial constraints while preserving perceived quality.
Improvements for AI systems
Here are the specific improvements to AI systems based on the GestAdapt framework, and what these improved systems can achieve:
The GestAdapt framework improves AI systems by fundamentally shifting co-speech gesture generation from a purely motion-centric task to one that is explicitly aware of and conditioned by the physical workspace constraints.
Here are the specific improvements and their resulting capabilities:
-
The introduction of a diffusion transformer model conditioned on speech, motion history, and a specified workspace box.
-
The integration of learned workspace boundaries (6 spatial tokens) into the denoising process via cross-attention mechanisms within the Transformer blocks.
-
The implementation of an analytic guidance term during sampling time to actively reduce residual wrist boundary violations in real-time generation without requiring additional training data or complex scene annotations.
-
The adoption of a unified motion representation (123D) that separates upper-body/palm variables (where workspace adaptation is primary) from finger articulation, using part-aware output heads to improve fidelity and efficiency.
-
The use of GMR (Generalized Motion Representation) for robot retargeting, which maps the generated abstract motion to target robot joint space using kinematic models, Jacobians, and joint limits while simultaneously applying soft workspace constraints via quadratic programming (QP).
These improved AI systems can achieve the following specific capabilities:
-
The system can generate a wide variety of plausible co-speech gestures that are not only semantically aligned with speech but are also physically feasible within a predefined spatial environment (e.g., next to a wall, over a table).
-
The system will produce motions that remain
close to the real-motion distribution
(high Frechet Gesture Distance) while strictly respecting the prescribed workspace boundaries, leading to high user satisfaction scores in real-world scenarios. -
The robot can dynamically adapt its gesture style based on its current physical context—for instance, switching from a broad arm swing in an open room to a constrained, localized movement when near an obstacle—without requiring the model to explicitly learn every possible spatial constraint beforehand.
-
The system will be highly effective for deployment on diverse robot embodiments (humanoid robots), as the GMR-based retargeting ensures that the generated motion is mapped accurately onto different kinematic structures while maintaining feasibility under strict joint limits and collision avoidance constraints.
-
The system can operate with high safety guarantees: by incorporating hard physical constraints (joint limits, self-collisions) during execution, it minimizes the risk of hardware damage or unsafe movements, a critical feature for autonomous robotics.
-
The system provides a superior solution to workspace adaptation compared to methods that only apply constraints after motion generation (downstream correction), as the learned conditioning directly influences the gesture synthesis process itself, resulting in demonstrably higher perceived quality by human evaluators.
Abstract
Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. Since the same speech can be accompanied by different gestures, a robot can respond to workspace constraints, e.g., gestures for speech next to a wall. In these scenarios, the robot should gesture in a suitable motion rather than simply correcting an unconstrained one. To achieve this goal, we present GestAdapt, a workspace-conditioned framework that conditions co-speech gesture generation on a prescribed wrist workspace. The GestAdapt framework learns from six complementary co-speech corpora through a shared motion representation and supports retargeting to different robot embodiments. Quantitative evaluation shows that generated motions remain close to the real-motion distribution while respecting the workspace. In a user study, gestures generated under modified workspace constraints receive a mean quality score of 3.24/5, above our no-workspace variant (2.43/5) and below the reference motions (3.68/5). In a real robot evaluation, all compared motions are retargeted to the Reachy2 humanoid robot under identical workspace constraints. Motions generated with our framework rank first in 69.7% of comparisons, higher than our no-workspace variant baseline and retargeted ground-truth motions constrained afterward. Overall, the results support adapting gestures to the available workspace during generation, rather than modifying unconstrained trajectories afterward to satisfy workspace constraints, potentially compromising gesture naturalness.
Sources
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving