VioLA: Learning Generalist Humanoid Control Policies from Human Data

arXiv:2610.12435 · cs.RO, cs.LG · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "VioLA: Learning Generalist Humanoid Control Policies from Human Data".

Dev: The gist The VioLA generalist humanoid policy learns to predict body and hand motion latents instead of joint commands,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're diving into VioLA today. It looks like this paper is tackling the big problem of teaching a humanoid robot to actually follow instructions for complex tasks without having to painstakingly fine-tune it for every single new thing it has to do.

Dev: Exactly, Rosa. The core idea here seems to be shifting away from trying to predict every single joint command directly, which is where most of the current humanoid policies get stuck because coordinating the whole body and keeping balance while reaching feels really hard for them <ref:2610.12435#pg2>.

Taro: And what I find interesting is how they bypass that coordination issue by focusing on predicting body and hand motion latents instead of joint commands, which lets the robot learn complex movements by predicting coordinated body and hand latents from visual input, instruction, and proprioception <ref:2610.12435#pg1>.

Rosa: So if I'm driving or just walking across the room, I might think, "Okay, how does this work in the real world? Does it actually handle the messy stuff?" We need to see if this works outside of a clean lab setting <ref:2610.12435#pg3>.

Dev: That's my main concern. The paper claims zero-shot performance on a real Unitree G1 for locomotion, achieving one hundred percent success, but then they drop the manipulation success rate to eighty-eight point six percent without any task-specific fine-tuning <ref:2610.12435#pg3>.

Taro: That distinction is important because it shows that while the generalist policy can handle basic movement instructions out of the box, getting the precise manipulation right still requires some kind of specific training for those fine motor skills.

Rosa: Right, so even though it can walk to a table and close a laptop without tuning it for that exact laptop closing motion, it struggles when we ask it to do something more intricate like hanging up a coat or rotating a chair <ref:2610.12435#pg3>.

Dev: The methodology they used seems pretty clever with this shared latent space where the body and hand latents are predicted together, so reaching and finger motion can be timed in sync, which is what lets the frozen controllers map those predictions into actual joint commands <ref:2610.12435#pg3>.

Taro: And they trained this policy using human demonstrations to supervise it without needing a corresponding physical robot demonstration for every task, creating seven hundred eighty-one hours of training data, with ninety-three point two percent of that being from humans <ref:2610.12435#pg3>.

Rosa: That reliance on human motion data is significant because it bypasses the massive bottleneck of getting enough demonstrations for every single task, which is something we’ve struggled with for a long time in humanoid research <ref:2610.12435#pg1>.

Dev: They tested this recipe across different backbones too, trying two VLA models and one WAM model, and they found that the GR00T backbone gave them the strongest results overall <ref:2610.12435#pg3>.

Taro: The finding that GR00T performs substantially better on manipulation trials than the other backbones suggests that for tasks requiring detailed interaction, selecting the right foundation model really matters for what you get out of the system <ref:2610.12435#pg3>.

Title and authors: Rosa: So when we look at these improvements they suggest, they’re focusing on how this latent space approach allows them to predict motions rather than just joint commands, which means the robot can learn complex movements by predicting coordinated body and hand latents from visual input, instruction, and proprioception Improvement one <ref:2610.12435#pg1>.

Dev: They also talked about using real-time chunking to enhance execution continuity, where the next prediction is conditioned on the latents scheduled to execute before inference finishes, which helps reduce discontinuities in the predicted body latents when consecutive chunks are joined Improvement five <ref:2610.12435#pg1>.

Taro: I’m interested in how they handled situations where things go wrong; they mention that failures occur at different stages for manipulation, showing that reaching the object or making an initial grasp isn't always enough for success <ref:2610.12435#pg3>.

Rosa: That limitation is important because it points toward where we need to focus next—getting the robot to actually retain an object once it's grasped and dealing with contact-force feedback, which they noted the hand decoder currently lacks Limitation one <ref:2610.12435#pg1>.

Dev: Speaking of limitations, they also mentioned that human supervision doesn't improve every behavior under their current training schedule; they found that robot-only training achieved higher manipulation success in some cases Limitation two <ref:2610.12435#pg1>.

Taro: That suggests there’s a trade-off between the breadth of general instruction following and the depth of specific, fine-grained control when we only have human data to work with Limitation two <ref:2610.12435#pg1>.

Rosa: So to wrap up on this VioLA paper, it shows that learning from human motion latents allows a generalist humanoid policy to achieve one hundred percent success on locomotion and about eighty-eight point six percent on manipulation without needing task-specific fine-tuning, even when tested on real hardware <ref:2610.12435#pg3>.

Dev: The main implication for engineers is that they can use human demonstrations to build a generalist policy that works out of the box for a wide range of movements, provided you select the right backbone like GR00T <ref:2610.12435#pg3>.

Taro: For autonomy researchers, it means we can start training models on vast amounts of human data to get that initial general capability, and then focus our efforts on what specific failures they encounter in complex manipulation tasks <ref:2610.12435#pg3>.

Rosa: It’s a step toward making humanoid deployment less dependent on massive task-specific datasets and more reliant on learning from the motions we already see humans do <ref:2610.12435#pg1>.

Dev: We need to keep an eye on those failure modes, especially regarding contact feedback and how it affects the robot’s ability to interact physically with objects Limitation one <ref:2610.12435#pg1>.

Taro: It's a big step in getting these systems to behave more like actual humans when they encounter unexpected situations <ref:2610.12435#pg3>.

Rosa: So that’s where we are on VioLA today. We’ll be back after the break to look at some other work on embodied AI Improvement six.

The paper's summary: Rosa: So, to wrap up what we just heard about VioLA, basically this paper is about training a humanoid policy that can follow instructions for almost any movement without having to teach it specifically for every single task.

Dev: Exactly. It moves away from trying to predict every single joint command directly and instead focuses on predicting body and hand motions as latents.

Rosa: That means the robot learns coordinated movements by looking at human examples, whether those examples are for walking or closing a laptop, without needing a separate fine-tuning session for each one.

Dev: It's that shared latent space where the body and hand predictions are linked together, so reaching and finger movements can be timed to each other automatically.

Rosa: And the numbers they showed are pretty telling—it hits one hundred percent success on basic locomotion when tested on a real robot, which is a big deal for real-world deployment.

Dev: But there's that caveat about manipulation; it gets eighty-eight point six percent success on tasks like closing a laptop without any task-specific fine-tuning, which suggests it’s great at general movement but needs some help with the fine motor details of grasping and manipulating objects.

Rosa: That points to where we need to keep looking—how does this system handle the messy parts of real interaction, like getting a grip or dealing with unexpected forces?

Dev: Right. And they did look into that, noting that the hand decoder doesn't have contact-force feedback yet, so it’s not fully equipped for those kinds of physical interactions.

Rosa: It shows that human demonstrations are a powerful way to supervise this kind of general policy, creating a lot of training data without needing tons of expensive robot demos for every new thing we want it to do.

Dev: And they found that the GR00T backbone was the strongest performer overall, matching other models on locomotion and doing better than them in manipulation tests.

Rosa: So what this means for us is that we can build a generalist base policy using human data, and then our real work shifts to figuring out how to give it better tools for those difficult physical interactions.

The paper's improvements: Rosa: So we’re looking at what these authors are suggesting for making VioLA even better, focusing on how they can push this generalist policy further Improvement one.

Dev: They’re talking about moving beyond just predicting joint commands and truly learning to predict those coordinated body and hand motions directly from visual input, instructions, and proprioception.

Rosa: It sounds like they want the AI to learn what the motion *is*, not just what every single joint needs to do for that motion Improvement one.

Dev: And they’re using human data as a direct supervision signal to train this policy, essentially turning human actions into targets so you don't need a specific robot demo for every new task Improvement two.

Rosa: That reliance on human motion helps build up that massive training set without having to painstakingly record every single robot action ourselves Improvement two.

Dev: They also talked about using real-time chunking during execution, which means the AI keeps predicting movements even while the current prediction is still finishing, which helps keep things smooth when it’s running fast Improvement five.

Rosa: That's smart. And they found that by picking the right backbone—like GR00T—you get much better performance on manipulation tasks compared to other options Improvement six.

Dev: So, the implication is that you don't need one single perfect model for everything; you can pick a foundation model based on what you want it to excel at, like locomotion or fine motor control Improvement six.

Rosa: But they did point out some sticking points too. They mentioned that the hand decoder doesn't give contact-force feedback, so it’s not great for those delicate physical interactions Limitation one.

Dev: And they also noted that just having human supervision doesn't always improve every single behavior; sometimes training purely on robot demonstrations actually gives better results for certain manipulation tasks Limitation two.

Rosa: So the big picture here is that VioLA sets a baseline for general humanoid control, but the next big step is giving it better physical interaction capabilities and figuring out how to balance instruction following with those fine motor constraints Limitation one Limitation two.

Conclusion: Rosa: So we've covered VioLA, learning generalist humanoid control policies from human data, and to wrap up, this system shows that training on human motion latents lets an AI learn complex movements without needing task-specific fine-tuning for every single thing it has to do.

Dev: That’s the main point; it’s a generalist policy trained on demonstrations that can actually follow instructions out of the box when deployed on a real robot.

Rosa: It really changes how we think about building these robots because we don't need millions of specific robot demos for every new thing, just human data to build the foundation.

Dev: Yeah, but we still have those practical hurdles like making sure the loop rates are tight and handling those failure modes when things go wrong in real-time.

Taro: I’m just thinking about how this generalist approach might let us focus our autonomy research on what happens when the world misbehaves, instead of spending all our time getting a policy to walk perfectly across the floor.

Rosa: That makes sense. And we still have those limitations, like it not having contact-force feedback for manipulation right now, which is a big gap for real-world interaction.

Dev: True; that lack of tactile sensing means the robot can predict what *should* happen but might fail when it actually touches something because it doesn't know the force involved.

Taro: It suggests that future work needs to heavily focus on incorporating contact feedback into hand control so this generalist foundation can become truly useful for complex physical tasks.

Rosa: So, VioLA is a solid step in using human data to build broad humanoid skills, and it opens up the door for us to explore those more nuanced physical interactions next.

Dev: Exactly; it’s a powerful starting point for embodied AI, but we still need to get those latency issues under control as we scale these policies up.

Taro: Next time, we should look at papers that tackle how these models handle visual bottlenecks when they are actually trying to perform precise actions.

Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius

Vesoma Institute of Technology Zurich Institute (Vesoma) · MPI-IS Institute

cs.RO, cs.LG

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://www.figure.ai/news/helix-02

The gist: The gist The VioLA generalist humanoid policy learns to predict body and hand motion latents instead of joint commands, enabling zero-shot locomotion and manipulation on a real robot without

Key concepts

Latent Action Space
This is a shared mathematical space where both body motion and hand motion predictions reside simultaneously. Instead of predicting specific joint angles, the policy predicts abstract 'latents' representing desired movements for the whole body and hands together. This allows the system to time reaching motions with finger movements correctly.
Generalist Policy
This is a single humanoid control policy trained on a broad range of human demonstrations. It is designed to be versatile, meaning it can perform various tasks like walking or closing a laptop without needing separate training for each specific action. It learns the underlying principles of human movement.
Zero-Shot Locomotion
This means the robot can successfully walk or move on a real physical robot immediately after being trained only on human demonstrations, without any further training tailored to that specific walking task. VioLA achieves 100% success in locomotion tasks using this method on a real Unitree G1.
Motion Latents
These are compressed representations of body and hand movements learned by the policy. Instead of controlling every individual joint, the policy outputs these compact latents. These latents are then fed into separate, pre-trained controllers that translate them into physical joint commands for the robot.

Terminology

Summary

The gist The VioLA generalist humanoid policy learns to predict body and hand motion latents instead of joint commands, enabling zero-shot locomotion and manipulation on a real robot without task-specific fine-tuning

How it works

VioLA introduces a generalist humanoid policy that predicts body and hand motion latents rather than joint commands, which are then executed by pretrained body- and hand-controllers. These encoders map human and robot motion into the same latent spaces, meaning a human recording is labeled in the policy’s action space. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot, reaching 100% success where GR00T N1.7 and Ψ0 reach 16.7% and 0%, respectively.

Shared Latent Action Space

To close a laptop, the robot must reach for the lid and push it down while remaining balanced. VioLA predicts body and hand latents together so that reaching and finger motion can be timed to one another. The action space is defined as zt = [zbody; zhand] where zbody is a body latent and zhand is a hand latent. This approach allows the generalist policy to predict motions, not joints, while frozen controllers map these predictions into joint commands.

Learning from Human and Robot Demonstrations

Human demonstrations supervise the generalist policy’s deployed action outputs without a corresponding physical robot demonstration. The human motion encoders turn recorded body and hand motion into action targets, and robot demonstrations enter the same space through the corresponding robot encoders. Both sources provide targets in the same action space, allowing them to train the same prediction head with the same loss. This method creates 781 hours of training data, 93.2% of which is human data.

Performance and Backbones

On a real Unitree G1, VioLA achieves 100% locomotion success and 88.6% manipulation success without task-specific fine-tuning for either, among closing a laptop, hanging up a coat, and rotating a chair. The approach is not tied to one model, and the same recipe is tested on two VLA and one WAM backbone. GR00T gives the strongest overall results: it matches DiT4DiT on locomotion and completes substantially more manipulation trials than either alternative.

Key Findings in Manipulation

VioLA succeeds in 31/35 manipulation trials (88.6%) without task-specific fine-tuning. However, failures occur at different stages, showing that reaching the object or making an initial grasp is insufficient for manipulation success. The robot-only policy completes 33/35 trials (94.3%), while VioLA completes 31/35 trials (88.6%).

Limitations and Future Directions

Limitations include the fact that the hand decoder produces finger-position commands without contact-force feedback, and remaining difficulties in reaching to grasp objects and retaining them through task completion. Incorporating contact feedback into hand control is one direction for improving object interaction. Human supervision also does not improve every behavior under the current training schedule, as robot-only training achieves higher manipulation success. The main evaluation covers thirteen tasks on one Unitree G1, with five trials per task, and broader testing across tasks, environments, and robot embodiments is needed.

The paper demonstrates that VioLA is a generalist humanoid policy that learns from human and robot demonstrations to predict body and hand movements. On a real Unitree G1, VioLA achieves 100% locomotion success and 88.6% manipulation success without task-specific fine-tuning. Its complex movements are represented with a few motion tokens. Comparing two VLA backbones and one WAM backbone using the same training data and budget shows that VioLA performs best with GR00T, particularly for manipulation. Training the generalist policy on human demonstrations alone still yields 40% locomotion success on the real robot. The gist.

The gist.

Improvements for AI systems

  1. Body and hand motion latent prediction allows for predicting motions, not joints, meaning The generalist policy selects body and hand motions rather than joint commands. This enables the robot to learn complex movements by predicting coordinated body and hand latents from visual input, instruction, and proprioception.

  2. Human demonstrations serve as direct supervision for the generalist policy by converting them into action targets: Human images, instructions, and encoded motion train the same generalist policy as robot demonstrations. This eliminates the need for expensive robot demonstrations for every task fine-tuning.

  3. The system achieves zero-shot performance on real hardware: VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning. This means the policy can execute novel instructions immediately upon deployment, reaching 100% success where GR00T N1.7 and Ψ0 reach 16.7% and 0%, respectively in locomotion tasks.

  4. The system utilizes a shared latent space for joint coordination: VioLA learns which body and hand motions an instruction requires while pretrained decoders produce the robot’s joint commands. This structure ensures that the frozen controller converts its predictions into joint commands, allowing for coordinated movement while maintaining balance.

  5. Real-time chunking (RTC) enhances execution continuity: The next prediction is conditioned on the latents scheduled to execute before inference finishes, which helps in reducing discontinuities in the predicted body latents when consecutive chunks are joined. This ensures smoother, more continuous motion during high-frequency control updates.

  6. Backbone selection optimizes performance based on task type: GR00T gives the strongest overall results: it matches DiT4DiT on locomotion and completes substantially more manipulation trials than either alternative. This allows researchers to select the most effective foundation model (VLA vs. WAM) for specific control objectives.

  7. Distinct motion tokens provide a compact representation of complex movements: Complex movements can be represented with a few motion tokens, such as a complete chair-rotation execution uses an average of 60 distinct body tokens and 31 hand tokens. This demonstrates the policy's ability to learn high-level, efficient control strategies.

Sources

Related papers