WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

arXiv:2607.11397 · cs.RO · Submitted 2026-07-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos".

Dev: WALA is a framework designed to jointly learn executable latent actions from both action-labeled demonstrations and action-free videos,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're diving into the paper "WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos," which tackles how to get good robot policies when labeled data is scarce, right?

Dev: Exactly, Rosa. The core idea seems to be combining the benefits of action-labeled demonstrations with the abundance of action-free human videos. It claims WALA can jointly learn these executable latent actions from both types of data.

Taro: I'm curious about what makes this joint learning approach so compelling for autonomy research, Rosa; does it really help when you only have a few robot examples?

Rosa: That's the question, Taro. The paper suggests that WALA first pretrains a semantic-geometric latent action model using videos that don't have robot action labels to learn representations from scene evolution. This allows it to ground those latent actions in actual physical changes in the scene, which is pretty key for generalization.

Dev: From an engineering standpoint, I'm interested in how this pretraining stage works without needing explicit robot control signals for those initial latent actions; we need to make sure the representations it learns are actually useful later on.

Taro: If it learns action-relevant representations from scene evolution, what happens when the world misbehaves during deployment? Does this representation hold up when things change unexpectedly outside of the training distribution?

Rosa: The paper focuses on learning latent actions that explain future semantic and geometric changes based on current observations and sparse future samples. This means it’s learning transitions rather than just static appearance details, which should give it some robustness in handling unexpected movement.

Dev: That focus on transitions is interesting because it ties the latent action directly to observable physical dynamics, which speaks to loop rate concerns; we want these actions to be predictable at high frequencies.

Taro: And what about the actual deployment? Rosa, you mentioned testing outside the lab; how long do you think this learned capability stays stable when a robot operates in an uncontrolled real-world environment?

Rosa: The paper includes real-robot studies, and they showed that WALA achieved a seventy-five point two percent average success rate on RoboCasa under certain settings, which suggests it can perform well even outside of highly controlled lab conditions.

Paper summary: Dev: Seventy-five point two percent is solid performance for a complex task; from a control engineering perspective, I'd want to know if the latency introduced by using this latent model at inference time is acceptable for real-time operation.

Taro: That brings up the inference part, doesn't it? If we look at how WALA works at inference time, it only needs the vision-language backbone and action head, without needing that heavy world model decoder running constantly.

Rosa: Exactly; that’s one of the key innovations mentioned in "WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos"; it avoids world model inference overhead during deployment.

Dev: That's a big win for latency, Rosa; removing the need to run a full world model prediction cycle at every step is crucial for real-time control loops. However, what happens if the latent action target matching loss doesn't align perfectly with those frozen targets?

Taro: If the alignment isn't perfect, I worry about catastrophic failure when the robot encounters a novel situation that deviates significantly from what was seen during pretraining.

Rosa: The paper addresses this by using three losses during policy training: robot action prediction, latent action target matching, and future dynamics prediction. This joint supervision is designed to constrain the latent actions effectively across those different objectives.

Dev: So the system isn't relying on just one source of supervision; it’s cross-checking the learned actions against both real robot demonstrations and the predictive model's understanding of future dynamics.

Taro: That combination sounds like a good way to handle misbehavior because it gives you multiple constraints on what constitutes a valid latent action under different scenarios.

Rosa: It seems WALA is designed to use that rich supervisory signal from both sources—the labeled data and the action-free videos—to build these constrained latent actions, which is the main claim of this paper.

Dev: And those constraints are what keep things stable when we move from simulation to reality, provided the learned representations are robust enough for those real-world scenarios.

Taro: Looking ahead, what kind of future work do the authors suggest? I'm thinking about how they might extend this beyond the specific manipulation benchmarks they tested.

Rosa: The paper points toward scaling up its use of action-free data and testing it on more diverse, complex manipulation tasks to see if that generalization holds across different physical scenarios.

Paper summary: Dev: If we could get a clearer picture of the exact computational cost during that pretraining phase, I'd want to see if it scales well when we move to even larger video inputs or higher frame rates for better dynamics capture.

Taro: That would be important because scaling up the input data often introduces new kinds of complexity in how those semantic and geometric deltas are calculated.

Rosa: So, while this paper focuses on establishing a joint learning framework, the implications are that we might not need perfectly annotated robot data to build strong policies if we can leverage video evolution effectively.

Dev: I agree; it shifts the focus from expensive labeling efforts toward leveraging existing human interaction data to provide useful dynamics supervision for robot control.

Taro: That really makes sense for the broader field of autonomy; if we can get decent performance with less labeled robot data, it opens up a lot more practical application possibilities.

Rosa: So, to wrap up on this paper "WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos," the main point is that it successfully learns executable latent actions by fusing labeled robot demonstrations with action-free video supervision through a two-stage learning process.

Dev: That fusion is what makes the method powerful, allowing it to learn representations grounded in both direct control examples and general physical dynamics captured in videos.

Taro: It's an interesting contribution because it shows that action-free human videos can provide meaningful dynamics supervision even when robot action labels are missing, which is a significant step for autonomy research.

Rosa: And the results, like the seventy-five point two percent success rate on RoboCasa, suggest that this approach can translate into competitive performance in real-world manipulation tasks.

Dev: From an engineering standpoint, the efficiency of using only the vision-language backbone at inference time is a major practical advantage that reduces overhead and potential latency issues.

Taro: I'm excited to see how researchers extend this idea to handle more complex, misbehaving environments where the learned latent actions still need to be adaptable.

Rosa: That's what we hope for; moving beyond the controlled benchmarks into truly unpredictable physical interactions is where the real test of WALA’s generalization will come.

Conclusion: Rosa: So we're wrapping up our discussion on WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos, which essentially shows how to teach robots to move by mixing labeled examples with general video data.

Dev: I agree, Rosa; the authors are doing something really interesting by jointly learning latent actions from both types of data.

Taro: What this means for autonomy is that we can build policies even when we don't have tons of specific robot control examples, which opens up a lot more practical application possibilities.

Rosa: Exactly, Taro; it suggests that the way robots move can be understood through scene evolution rather than just looking at specific commands.

Dev: From an engineering standpoint, that fusion of knowledge is what makes the method powerful for creating stable control signals, even when things get messy in a real-world setting.

Taro: I'm thinking about the long-term impact; if this works reliably outside of a controlled lab environment, it could significantly lower the barrier to entry for deploying complex robotic systems in unstructured settings.

Rosa: That’s what we’re hoping to see, Taro; the results on diverse manipulation tasks show some real promise for that kind of generalizability.

Dev: And if the authors can keep that inference overhead low while maintaining good control, then we could see this kind of latent action approach being used in more demanding applications very soon.

Jiahao Liu, Zhongpu Xia, Shuai Tian, Huangrui Li, Yuhang Zheng, Ning Ma, Xin Fu

CASIA · Anyverse Dynamics

cs.RO

Submitted: 2026-07-13

Updated: 2026-09-28

Comments: Project page: https://liujiahao2077.github.io/WALA.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: WALA is a framework designed to jointly learn executable latent actions from both action-labeled demonstrations and action-free videos, addressing the limitations of current vision-language models by

Key concepts

Semantic and Geometric Future Deltas
Instead of predicting raw pixels, WALA predicts how the scene's features (semantic content and 3D structure) will change between the current frame and several future frames. Semantic deltas capture task-relevant object or state changes, while geometric deltas describe spatial structure shifts in depth space. This focuses learning on meaningful transitions rather than static visual details.
Pretraining Stage
The initial phase trains a latent action model using videos that lack robot action labels. It uses the current frame and future frames to learn what latent actions cause specific semantic and geometric changes. This stage forces the model to focus on 'action-induced transitions' by optimizing losses related to feature deltas and depth regression, ensuring the learned representations are task-relevant.
Policy Training Stage
In this stage, the pretrained encoder is frozen for stability while a trainable decoder acts as a latent world model. The policy learns unified latent actions guided by robot states and language instructions. These actions are supervised by three losses: matching the learned targets, predicting future dynamics via the world model, and ensuring executable robot control.
Action-Free Data Utility
WALA demonstrates that action-free videos can provide valuable supervision for real-world control. By using these videos to train the future dynamics prediction component of the policy, WALA achieves strong performance on manipulation tasks even when robot demonstrations are scarce. This shows that observing scene evolution is crucial for grounding actions in real-world scenarios.

Terminology

Summary

WALA is a framework designed to jointly learn executable latent actions from both action-labeled demonstrations and action-free videos, addressing the limitations of current vision-language models by leveraging future scene evolution to ground latent actions in task-relevant semantic and geometric changes.

The gist: WALA first pretrains a semantic-geometric latent action model on videos without action annotations, enabling it to learn action-relevant representations from scene evolution between the current observation and multiple sparsely sampled future observations.

How it works (Pretraining Stage)

WALA employs a two-stage process where the first stage involves pretraining a semantic-geometric latent action model on videos lacking robot action labels. This model uses the current frame and future deltas computed from multiple sparsely sampled frames to learn latent actions that explain future semantic and geometric changes. Specifically, the framework forms semantic and geometric future deltas, from which the encoder extracts latent action targets. The encoder captures semantic changes in DINOv3 feature space, while a separate branch captures geometric changes in dense depth space. This is achieved by using a frozen DINOv3 encoder to extract features and obtaining dense depth maps from depth observations or a frozen depth estimator. The objective function for this stage is defined as:

LLAM = Lrgb + λdepLdep, where Lrgb uses an l1 term and a cosine term between predicted and target DINOv3 feature deltas, and Ldep uses a dense depth regression term and a depth gradient consistency term. This process encourages the latent actions to focus on action-induced transitions rather than static appearance details.

How it works (Policy Training Stage)

In the second stage, WALA integrates this pretrained model into policy learning. The pretrained encoder remains frozen to provide stable latent action targets, while the decoder is used as a trainable latent world model. The vision-language backbone generates unified latent actions from multi-view observations, language instructions, robot states, and action queries. These generated latent actions are jointly supervised by three components:

  1. Robot action prediction (Lact): This loss ties the latent actions to executable control.

  2. Latent action target matching (Lalign): This loss ensures the generated actions are aligned with the frozen latent action targets.

  3. Future dynamics prediction (Lwm): The decoder is used to predict future semantic and geometric deltas, which keeps the latent actions grounded in future scene evolution.

The final policy training objective is:

Lpolicy = mLact + λalignLalign + λwmLwm. For action-labeled demonstrations, all three losses are used; for action-free videos, the action loss is masked out, while latent action target matching and future dynamics prediction remain active. This joint learning scheme ensures the latent actions are constrained by observed future scene evolution and tied to robot actions through action-labeled demonstrations.

Key Components and Innovations

The core innovation lies in how WALA models future changes. Instead of reconstructing raw pixels or absolute future observations, WALA predicts future deltas in the DINOv3 feature space and dense depth space. The semantic deltas capture task-relevant object and state changes, while the depth deltas provide geometric supervision over spatial structure. This avoids raw pixel reconstruction, reducing the influence of appearance details while preserving task-relevant semantic and geometric structure. Furthermore, at inference time, WALA is designed to be efficient; it uses only the vision-language backbone and action head, meaning future observations, the frozen latent action encoder, the DINOv3 encoder, the depth estimator, and the latent world model decoder are not required, thus avoiding world-model inference overhead.

Performance and Real-World Validation

Experiments on simulated benchmarks show strong performance. On RoboTwin 2.0 under Random settings, WALA achieved a 92.8% success rate, which was competitive with or better than existing VLA-based and WAM-based methods. On RoboCasa-GR1-Tabletop, WALA reached an average success rate of 75.2%, setting a new state-of-the-art result by outperforming the strongest reported baseline (DIAL) by 5.0 percentage points. Scaling studies confirm the utility of action-free data: adding action-free RoboCasa videos to a fixed set of labeled demonstrations further improves performance, showing that videos without robot action labels can still provide useful dynamics supervision. Real-world experiments validate this, where WALA achieved a 75.0% average success rate on four diverse manipulation tasks when combined with 400 action-free human videos in a low-label setting, nearly matching the performance of policies trained with more robot demonstrations. This demonstrates that action-free human videos can partially compensate for limited real-robot demonstrations by providing useful dynamics supervision for real-world control.

Contributions

The main contributions are summarized as follows:

Improvements for AI systems

Here are specific improvements that could be made to existing AI systems based on the WALA framework, along with what those improved systems could achieve:


  1. The ability of robotic policies to generalize across diverse object appearances and scene layouts without extensive task-specific retraining (as demonstrated by the zero-shot transfer capability).

  2. The capacity for vision-language models to learn a robust, executable latent action space that is inherently grounded in physical scene geometry and semantic evolution, rather than just mimicking low-level pixel changes.

  3. The creation of highly data-efficient robot control policies that can effectively leverage large volumes of readily available, action-free human videos (e.g., 400+ videos per task) to provide superior dynamics supervision compared to purely action-labeled datasets.

  4. The development of future dynamics prediction as a universal training signal that bridges the gap between static observation and executable control, allowing policies to understand the consequences of an action on the future state of the world before it is taken.

  5. The deployment of inference-efficient robot control systems where only a vision-language backbone and an action head are required at runtime, eliminating the need to run computationally expensive latent world models or complex encoders during live operation.

These improvements can lead to the following specific capabilities for AI systems:

  1. A robot that can successfully perform a novel manipulation task (e.g., grasping bread and placing it on a dinner plate) immediately upon being deployed, despite never having been explicitly trained on that exact combination of object and scene layout.

  2. A vision-language agent capable of interpreting natural language instructions (Put the mug on the shelf) and translating that into a sequence of latent actions that are guaranteed to result in a physically plausible future state change (semantic and geometric consistency).

  3. A robotic arm or humanoid system that achieves high success rates in real-world, out-of-distribution scenarios (e.g., grasping an unfamiliar object) by leveraging the general physical interaction knowledge embedded in vast action-free human videos, effectively reducing the need for millions of expensive robot demonstrations.

  4. A policy learning pipeline where the system learns what to do next not just based on what it sees now, but on a prediction of how that action will change the scene (e.g., predicting that grasping an object will cause it to move or change its position in space).

  5. A low-latency robotic control system suitable for real-time applications (e.g., manufacturing or household assistance) where the complex world modeling components are confined to the training phase, ensuring fast, deployable inference speed on edge hardware.

Sources

Related papers