SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation".
Dev: World action models (WAMs) combine robot action generation with future state prediction,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: Thinking about the title, SkeleWAM: Skeleton World-Action Modeling for Efficient Robotic Manipulation, I see that the authors are focusing heavily on creating a compact WAM specifically by using this sparse three dee skeleton representation instead of traditional implicit methods like videos or latent features. This suggests their main contribution is making the state space explicit and geometrically meaningful for action learning.
Dev: I agree with Rosa; it moves away from relying on visual reconstruction to parameterizing geometry, which should definitely simplify the architecture and improve inference speed, even if we have to be careful about how much information gets lost in that sparsification process. The authors claim this explicit geometric state achieves a favorable trade-off among task success, inference speed, model size, and computational cost.
Taro: From an autonomy perspective, the implication is that we are building models that don't just learn to mimic observed movements in video space but instead learn a representation of the physical world's structure itself, which is a much more fundamental way for an AI to understand manipulation. This structural understanding should make it more robust when the environment changes slightly.
Rosa: It sounds like SkeleWAM suggests that for complex manipulation tasks, focusing on the underlying kinematic and interaction geometry provides a very effective state space that is both efficient and rich enough for action learning without needing massive visual inputs. I'm wondering how this explicit structure performs when we move beyond the benchmark settings into truly unstructured environments.
Dev: That's where our concerns about failure modes come back in, Rosa; if the model relies so heavily on this specific geometric tokenization, what happens if the perception network used to create those initial object centers or interaction points gives us noisy data? The stability of that learned skeleton representation under sensor noise is a real question for me.
Taro: If we can design future work that allows the skeleton itself to be dynamically refined during training based on interaction feedback, that could solve some of those representation issues you're worried about, Rosa. Learning the representations beneficial for action generation through future skeleton supervision during training was something they highlighted as crucial.
Rosa: So, in simple terms, SkeleWAM is a compact world action model that uses a sparse three dee skeleton of robot joints and object points to jointly predict actions and future scene states, offering efficiency by avoiding complex visual predictions. The authors are pointing toward the value of explicit geometry for state representation.
Dev: And for the engineering side, it means we need to ensure that this explicit geometric tokenization doesn't introduce unacceptable latency when running in a high-frequency control loop, which is something I'll keep tracking as they move toward real deployment.
Taro: It gives me a feeling that this work sets a good foundation for future autonomy research by showing that we can build powerful world models using structured, geometric representations instead of just dense visual ones.
Rosa: It’s certainly an exciting direction to look into, and I'm eager to see how the team continues to push these ideas forward in handling real-world unpredictability.
Conclusion: Rosa: So, to wrap up this part of our discussion, SkeleWAM is essentially proposing a method where we build the world action model by representing the scene as a sparse three dee skeleton made of key points and object locations, which allows the AI to learn both what action to take and what the future scene will look like. Dev, when you look at that title and those authors, how do you see this approach fitting into the broader picture of robotic control?
Dev: I see it as a major step toward making these models more compact because they’re not trying to process dense visual data; they're using this geometric skeleton instead, which should naturally lower the computational overhead for real-time use. The authors are focusing on achieving efficiency while still maintaining a strong link between the scene structure and the learned actions.
Taro: Exactly, and from an autonomy standpoint, if we can decouple action generation from needing to constantly re-process high-dimensional images, that opens up possibilities for more reactive systems where understanding physical relationships is key. I’m thinking about how this explicit geometry helps when things in the environment don't behave exactly as expected.
Rosa: That makes sense, Taro, and the authors seem very confident about this structural representation; they claim it provides a much better state space for learning than previous implicit methods did. But I have to ask, Dev, where does that explicit geometric representation hold up when we take this out of the controlled lab setting and throw it into a messy real-world environment? How long can we really expect these models to function reliably there?
Dev: That’s my main concern, Rosa; the stability of those learned skeleton features under real-world sensor noise is what we need to test. We need to know if this model can handle the inherent unpredictability of physical interaction outside a perfect simulation environment. The success rate on the benchmark was good, but that doesn't tell us much about robustness in unstructured settings yet.
Taro: If it does struggle with real-world noise, maybe the next step for this research is to build mechanisms that allow the skeleton itself to adapt or refine based on immediate feedback from unexpected physical events, rather than relying solely on a fixed initial structure. That kind of dynamic learning would be important for true autonomy.
Rosa: So we're looking at a very efficient model with strong geometric foundations but still needing more proof on its endurance in messy reality, and that leads us right into how this work might eventually impact the way we design robots for complex tasks.
Juyi Sheng, Hua Wang, Mengyuan Liu
Peking University
cs.RO
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://skelewam-project.github.io
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 91/100
The gist: World action models (WAMs) combine robot action generation with future state prediction, and SkeleWAM introduces a compact WAM that represents manipulation scenes as sparse 3D skeletons composed of
Key concepts
- Skeleton World Representation
- This represents a manipulation scene as a sparse 3D skeleton made of robot joint positions and key object landmarks. It captures the essential geometric structure relevant to the task, making it easier for the model to understand where things are in space.
- Skeleton World Action Model
- This is a two-expert Mixture-of-Transformers design that takes skeleton tokens as input. One expert predicts action vector fields, and another predicts future skeleton states. They work together to learn how robot movements relate to the scene's geometry.
- Medoid Action Consensus (MAC)
- During inference, MAC selects a robust action by comparing multiple candidate trajectories. It looks at the first few steps of several potential paths and chooses the one that is most representative, improving reliability in action selection.
Terminology
Summary
World action models (WAMs) combine robot action generation with future state prediction, and SkeleWAM introduces a compact WAM that represents manipulation scenes as sparse 3D skeletons composed of robot joints, object centers, and interaction points. This approach is significant because it provides an explicit geometric representation of the relevant scene structure for action learning while substantially reducing model complexity compared to methods that rely on implicit visual representations like videos or latent features.
The gist
SkeleWAM represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points, jointly learning robot action generation and future skeleton prediction.
Problem Formulation
The objective is to generate an action chunk at time step t given a language instruction l, the current RGB-D observation ot, and the robot proprioceptive state qt. SkeleWAM represents the current scene as a sparse 3D skeleton, St = Φψ(ot, qt, l), which is encoded as a set of geometric tokens.
For training, the model jointly learns robot action generation and future skeleton prediction:
-
Action generation: p a θ(At St, l)
-
Future skeleton prediction: p s θ(S+t St, l)
Skeleton World Representation
The scene skeleton is defined by keypoints and task-relevant object landmarks. For each object m, the representation includes its center (ct,m) and a set of interaction points (ut,m,j). The complete skeleton is formed by concatenating robot joint positions (Jt) with the object landmarks: St = [Jt; Pt,1;...; Pt,M]. Robot nodes follow kinematic connectivity. Robot keypoints are computed via forward kinematics, and object landmarks are estimated from RGB-D observations using a fixed perception network during training.
Skeleton World Action Model
SkeleWAM employs a two-expert Mixture-of-Transformers design:
-
The
World expert
processes current and future skeleton tokens with shared parameters. -
The
Action expert
processes action tokens, both conditioned on the language instruction Cl = Elang(l).
The input is tokenized: Xc t = ϕc(St), Xa t,τa = ϕa(At,τa), and Xs t,τs = ϕs(Se+t,τs). The model predicts vector fields vˆ a θ ∈ R Ha×da and vˆ s θ ∈ R Hs×N×3 in a single forward pass.
Training Objective
The joint objective is defined using flow matching to supervise both prediction tasks:
-
Action loss: Laction = E[MSE (vˆ a θ, v a)], where v a = ϵ a − At.
-
Skeleton loss: Lskeleton = E[MSE (vˆ s θ, v s)], where v s = ϵ s − S+t.
The total loss is L = Laction + λsLskeleton, where λs weights the skeleton supervision during training.
Inference and Medoid Action Consensus (MAC)
At inference, future skeleton prediction is omitted. Actions are generated by integrating the learned action vector field: dA t dτ = vˆ a θ(Ae t,τ, St, l, τ: 1 → 0), yielding the sampled action chunk At = Ae t,0. To improve sampling robustness against stochasticity, SkeleWAM uses Medoid Action Consensus (MAC). MAC selects a representative trajectory from K ≥ 2 candidates by comparing their first h steps over dm continuous motion dimensions in normalized action space: ci = (1/K − 1) Σ X j≠i d ij, and the selected action is AMAC t = A(isel).
Evaluation
SkeleWAM was evaluated on the full LIBERO-Plus benchmark. In primary observation-based settings, it achieved an overall success rate of 85.9% with 57.1M parameters, outperforming CosmosPolicy by 3.7 percentage points and achieving 93.4% success under camera perturbations (sim-state variant). Real-world experiments on the ARX R5 robot showed SkeleWAM achieving an average success rate of 89% across five tasks, exceeding competitors like CosmosPolicy and Fast-WAM. Ablation studies confirmed that combining object centers with interaction points yields better results, and future skeleton supervision during training is crucial for learning representations beneficial for action generation. The final model uses He = 16 actions before replanning, and h = 10 for the MAC consensus window.
Conclusion
SkeleWAM demonstrates that explicit robot–object geometry provides an effective and parameter-efficient state space for world action learning.
Improvements for AI systems
Here are the specific improvements that an AI system could make by implementing SkeleWAM, along with a description of what the improved system can achieve:
The implementation of SkeleWAM enables a fundamental shift in how robotic manipulation tasks are modeled, moving from implicit visual or dense dynamic representations to an explicit geometric state space.
Here are the specific improvements and capabilities:
-
"""
-
Representation Shift: Replace complex, high-dimensional visual latents or dense 3D video frames as the primary world state input with a compact, sparse 3D skeleton representation (robot joints + object centers + interaction points).
-
Geometric Grounding for Control: Utilize the explicit geometry of robot-object interactions (e.g., contact points, relative joint positions) directly as the state representation for action generation, rather than relying on indirect visual features.
-
Joint Action and Future State Learning: Train a single model to simultaneously predict future geometric states (skeleton prediction) and generate immediate actions from the current skeleton state, using future prediction as auxiliary supervision during training.
-
Parameter Efficiency: Achieve high performance (e.g., 85%+ success rate on LIBERO-Plus) with significantly fewer parameters (e.g., 57M for SkeleWAM vs. billions for VLA/Video models), leading to faster training, lower inference costs, and deployment on resource-constrained hardware.
-
Robustness to Appearance Changes: The model is trained on geometric structures and learns to generalize actions based on this structure, making it more robust to variations in lighting, background clutter, and camera viewpoint compared to purely visual models.
-
Efficient Inference via MAC: Employ the Medoid Action Consensus (MAC) strategy during inference. Instead of relying solely on a single stochastic sample or complex reward models, MAC selects the most representative trajectory chunk from several noisy samples based on geometric distance in action space, improving sampling robustness without requiring a reward model or extensive averaging.
-
Unified State Representation: The system maintains a consistent robot-centric frame and normalization scheme for all nodes (joints and object landmarks), ensuring that the learned state is geometrically meaningful regardless of the specific scene configuration.
The improved AI system (SkeleWAM) can perform the following specific tasks:
-
"""
-
Zero-Shot Manipulation: Execute complex manipulation tasks (like opening a drawer, stacking bowls, or putting a block in a drawer) successfully without requiring task-specific fine-tuning on the evaluation dataset (LIBERO-Plus).
-
Real-Time Control Execution: Generate and execute continuous sequences of robot actions directly from the current sparse skeleton state and language instruction during inference (using flow integration), enabling dynamic, responsive manipulation.
-
Generalist Manipulation: Apply learned skills across diverse environments and object configurations due to its explicit geometric state representation, generalizing well beyond the specific scenes seen during pre-training.
-
Efficient Policy Deployment: Deploy a highly compact model (57M parameters) that offers competitive performance against large VLA models but with much lower computational overhead for real-time robotics applications.
Abstract
World action models (WAMs) combine robot action generation with future state prediction. Existing WAMs typically predict videos or learned visual latents, which represent interaction geometry only implicitly and may retain appearance information unrelated to control. We introduce SkeleWAM, a compact WAM that represents a manipulation scene as a sparse 3D skeleton composed of robot joints, object centers, and interaction points. Constructed online from current RGB-D observations and robot proprioception, the skeleton provides a unified geometric state for action generation and future skeleton prediction. Future skeleton prediction provides additional geometric supervision for action learning without requiring visual reconstruction. At inference, SkeleWAM generates actions directly from the current skeleton and language instruction, while Medoid Action Consensus (MAC) serves as an auxiliary consensus strategy for stochastic action samples. On LIBERO-Plus, SkeleWAM achieves an overall success rate of 85.9% with 57.1M parameters, outperforming Cosmos-Policy by 3.7 percentage points. These results demonstrate that sparse 3D robot--object structure provides an effective state space for robust and parameter-efficient world action learning. The project is available at https://skelewam-project.github.io/.
Sources
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- World Action Models are Zero-shot Policies
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving