SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills

arXiv:2610.12046 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills".

Dev: The gist: SkillWeave introduces a framework that learns from multi-modal demonstration by combining teleoperation for long-horizon coverage and kinesthetic teaching for precise,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're looking at this paper called "SkillWeave: Weaving Heterogeneous Demonstrations into Long-Horizon Manipulation Skills." Basically, they’ve got a problem where you need both big movements and really careful, precise touches to do a complex task, and collecting demonstrations that cover both styles of interaction is tough.

Dev: Right. They propose this framework that learns from two different kinds of demonstrations at once: teleoperation for the long-range stuff, like moving things around the workspace, and kinesthetic teaching for those fine, contact-rich skills where you need a lot of precision.

Taro: So it sounds like they’re trying to build one system that handles both coarse reaching and detailed interaction effectively. But collecting data for those two regimes usually requires different kinds of demonstrations, which is the core difficulty here.

Rosa: Exactly. The key claim in SkillWeave is that combining these two modalities into one heterogeneous framework lets them achieve a twenty-seven percent average end-to-end success rate across three different real-world long-horizon tasks <ref:2610.12046#pg3>.

Dev: That's a solid number, but the mechanism they use to make it work is pretty clever. They address the visual mismatch that happens when you collect kinesthetic data because the human demonstrator’s hand gets in the way during training, but not when the robot does it autonomously.

Taro: How do they handle that visual mismatch? Because if you just remove the demonstrator, you lose information about how to interact precisely, which sounds like a big loss.

Rosa: They don't remove it or paint over it; instead, they propose something called an object-mask-conditioned diffusion policy. This means during training, they use offline object segmentation to generate masks for the objects involved in the interaction, and at deployment, they have a learned mask predictor that helps condition the policy using those masks.

Dev: So during training you use SAM three to get those offline masks for supervision, but when it's running autonomously, it uses this lighter learned predictor instead of something like image inpainting or SAM three directly on the control loop <ref:2610.12046#pg3>. That keeps things moving faster.

Taro: I wonder if relying on these masks works well when the object moves or changes shape during execution, though they seem to assume the task-relevant objects remain observable for this approach to work.

Paper summary: Rosa: That's a fair question, Taro. The paper points out that methods relying on visible embodiments have limitations because they need calibrated robot geometry and test-time masking, and their masks can still obscure things near contact regions. So SkillWeave is designed around object-centric methods that suppress irrelevant appearance through structured representations instead of trying to reconstruct every pixel.

Dev: And for the policy itself, they break down each sub-task into a diffusion model conditioned on visual features from both an external and a wrist camera, plus proprioception data like end-effector pose and hand pose. That feeds into a one-dimensional U-Net that predicts action chunks.

Taro: So for the control side, it’s using these visual features combined with the object masks to guide a diffusion model that outputs action chunks, which sounds like they're conditioning an AI policy on what it *should* be seeing based on where the object is.

Rosa: Right. Now for the composition part. The paper acknowledges that training each sub-task independently creates a problem because of this terminal-to-initial distribution mismatch between policies.

Dev: They call this a known limitation of naive skill composition, where small prediction errors in one step cause the next policy to start from a state that its demonstrations didn't cover, which can lead to failure even if the individual tasks succeed.

Taro: So how do they fix that distribution mismatch? They introduce something called successor-aware terminal steering. This is where the predecessor policy samples actions, checks where it lands based on the successor's demonstrated starting states, and steers toward those better states before handing off control.

Rosa: That steering mechanism repeatedly samples action chunks from one policy to guide the system toward a state supported by the next one’s demonstration distribution. This is what allows them to compose policies that are independently trained but work reliably in sequence.

Dev: And they show that this composition efficiency is quite good, achieving an average of eighty-seven percent for long-horizon composition, which shows how much that steering actually helps compared to just running the policies sequentially without it.

Taro: What does this mean for autonomy? If you're doing a sequence of actions, and one step drifts off because of accumulated error, this steering process acts like a corrective guide to pull you back onto the path supported by the next demonstrated skill.

Paper summary: Rosa: It means that instead of needing separate transition demonstrations or a whole new policy just to bridge the gap between steps, SkillWeave allows you to use whatever demonstration modality is best for each step and stitch them together reliably using this steering.

Dev: So, in simple terms, they’ve combined teleoperation for coarse movement and kinesthetic teaching for fine control, used mask conditioning to handle the visual differences during learning, and implemented successor-aware steering to fix the handoff problems between those learned pieces.

Taro: It shows that matching the demonstration modality to the specific interaction regime is key—teleoperation for reaching, kinesthetic teaching for precision—and then using that steering mechanism fixes the sequencing issue.

Rosa: The implication here is that for long-horizon manipulation, you don't need one perfect demonstration type; you can use the best tool for each part of the job and use this steering to make it all flow together.

Dev: The caveat they mention is that even with this steering, if the prediction errors accumulate too much across a long sequence, that mismatch can still compound and cause failure in deployment.

Taro: So, while SkillWeave gets them to twenty-seven percent success on those tasks like cube reorientation or nut removal, the system isn't completely immune to errors just because it’s composed this way <ref:2610.12046#pg3>.

Rosa: That's right. The paper shows that combining these specific learning techniques—heterogeneous demonstrations and successor-aware steering—is a way to build reliable long-horizon behaviors without needing extra transition data or training a separate policy for the next step.

Dev: It’s about making sure the system doesn't just succeed at small steps, but actually stays on track as it moves through a whole complex task.

Taro: So for someone listening who only cares about what this means practically, it suggests that complex physical tasks don't have to be learned all at once with one massive dataset; you can break them down and use the right training style for each piece, then use a smarter composition strategy to put them back together.

Rosa: That's the big picture of SkillWeave: using complementary strengths—teleoperation and kinesthetic teaching—and a smart handoff mechanism to make long-horizon manipulation more robust.

Conclusion: Rosa: So, SkillWeave is this new framework that tries to bridge the gap between teleoperation for big movements and kinesthetic teaching for precise contact skills in one system.

Dev: Yeah, it’s about learning from different kinds of demonstrations—coarse reaching and fine manipulation—to get better results on those long-horizon tasks.

Taro: What I find interesting is how they handle that visual mismatch between collecting data from a human demonstrator and having the robot actually execute the skill.

Rosa: They solve that by using something called an object-mask-conditioned diffusion policy, which uses offline segmentation during training and a lightweight predictor for when it's running autonomously.

Dev: That’s smart because it keeps the online control loop from getting bogged down in complex visual reconstruction or inpainting.

Taro: And they didn't just stop there; they addressed the problem of chaining these separate skills together by introducing successor-aware terminal steering.

Rosa: So, instead of just stitching policies together and hoping for the best, this steering mechanism actively guides the system toward a state that matches what the next skill is supposed to start from.

Dev: That’s where you see a real improvement in composition efficiency; they report an average of eighty-seven percent for long-horizon composition compared to just running them separately.

Taro: What this means for autonomy is that you can learn different parts of a complex task using the best training data available for that part, and then use this steering to reliably string them together.

Rosa: It suggests that instead of needing one massive demonstration set covering everything, you can use specialized learning techniques and composition strategies to build reliable sequences.

Dev: But we have to keep an eye on those long chains; the paper does mention that even with steering, if errors accumulate over many steps, the whole sequence can still drift off.

Taro: Exactly. It’s a way to make the system more robust against small errors in each step during a long task.

Rosa: So, SkillWeave seems to be showing how combining these different learning approaches with that smart handoff mechanism is making complex physical tasks much more manageable for AI systems.

Dev: It’s moving the goal from just mastering single skills to actually performing long sequences reliably in the real world.

Ryosei Tamura, Xiaoxiang Dong, Uksang Yoo, Yuemin Mao, Romina Mir, Jonathan Francis, Jeffrey Ichnowski

Robotics Institute, Carnegie Mellon University, USA

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist: SkillWeave introduces a framework that learns from multi-modal demonstration by combining teleoperation for long-horizon coverage and kinesthetic teaching for precise, contact-rich

Key concepts

Teleoperation
This involves controlling the robot remotely, typically for coarse movements and transport. In SkillWeave, teleoperated demonstrations are used because they have little visual mismatch between the demonstration and the final autonomous execution, making them suitable for learning large-scale motion.
Kinesthetic Teaching
This method involves direct physical interaction where a human guides the robot's arm to learn precise, contact-rich skills. This is used when fine dexterity and accurate manipulation are required. The framework addresses the visual mismatch during this data collection by using object masks.
Object-Mask-Conditioned Diffusion Policy
This mechanism is used to train policies from kinesthetic demonstrations while accounting for the demonstrator's presence. It uses offline object segmentation to generate masks and a lightweight predictor at deployment, allowing the policy to execute autonomously without needing real-time image inpainting or SAM 3.
Successor-Aware Terminal Steering
This is a closed-loop control procedure that connects independently learned sub-task policies. It guides the system by repeatedly sampling actions from the current policy and checking if the resulting state aligns with what the next policy is expected to handle, ensuring smooth handoffs between skills.

Terminology

Summary

The gist: SkillWeave introduces a framework that learns from multi-modal demonstration by combining teleoperation for long-horizon coverage and kinesthetic teaching for precise, contact-rich interaction to achieve 27% average end-to-end success across three real-world long-horizon tasks.

How it works

SkillWeave is a heterogeneous demonstration framework that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills The approach addresses the challenge of collecting demonstrations that effectively support both regimes To handle the visual mismatch introduced by the demonstrator’s presence during kinesthetic data collection, the framework proposes an object-mask-conditioned diffusion policy This mechanism uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment This removes SAM 3 and image inpainting from the online control loop

Sub-Task Policies

For each sub-task T(i), the framework trains a diffusion policy π(i) from its corresponding demonstration dataset D(i) The observation space consists of ot = Iextt, Iwristt, qt, where Iextt and Iwristt are RGB observations from a thirdperson workspace camera and a wrist-mounted camera facing the robot palm Each camera observation is encoded by a ResNet-18 visual encoder ϕv The resulting visual features are combined with qt to condition a 1D U-Net ϵθ that denoises an action chunk

Teleoperation and Kinesthetic Policies

Teleoperated demonstrations contain only the robot and objects in the workspace and therefore exhibit little visual mismatch between demonstration and deployment Mask-Conditioned Kinesthetic Policies introduce a different observation distribution because the demonstrator’s hand and arm are visible during data collection but absent during autonomous execution Rather than removing or inpainting the demonstrator, the framework retains the original RGB observations and conditions visual feature aggregation on a mask of the manipulated object During training, object masks are generated offline using SAM 3 for each camera view

Policy Composition and Successor-Aware Handoff

The framework addresses the terminal-to-initial distribution mismatch between independently trained sub-task policies by introducing successoraware terminal steering This steering mechanism selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor’s demonstrated initial-state distribution The procedure involves repeatedly sampling candidate action chunks from π(i), evaluating their predicted terminal states according to proximity to the successor’s demonstrated start-state distribution, and executing the best candidate This closed-loop steering procedure moves the system toward a state supported by the successor’s training distribution, after which control is handed off to π(i+1)

Key Findings

The results show that matching demonstration modality to interaction regime improves long-horizon, contact-rich manipulation Teleoperation is well-suited for coarse arm-level motions, whereas kinesthetic teaching provides an advantage for skills requiring precise dexterous control Mask conditioning enables policies trained from kinesthetic demonstrations to execute autonomously despite the absence of the demonstrator SkillWeave achieves end-to-end success rates of 45%, 15%, and 20% on Cube Reorientation, Nut Removal and Storage, and Kettle Preparation, respectively, corresponding to an average success rate of 27% Successor-aware terminal steering mitigates the degradation in end-to-end success caused by naively composing the same sub-task policies without steering The composition efficiency results in table III further demonstrate the importance of terminal steering, achieving an average of 87% for long-horizon composition

Conclusion

SkillWeave successfully exploits the complementary strengths of teleoperation and kinesthetic teaching, using successor-aware terminal steering to compose independently learned policies into reliable long-horizon behaviors Future work could automate task decomposition and demonstration modalities decisions by jointly learning these from data The framework demonstrates that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged. The framework enables each sub-task to be learned using a demonstration modality suited to its control requirements and reliably composed into long-horizon behaviors The approach requires neither additional transition demonstrations nor a separately trained transition policy, while leaving π(i+1) unchanged.

Improvements for AI systems

  1. The system can achieve long-horizon dexterous manipulation by combining teleoperation for coarse arm motions with kinesthetic teaching for precise, contact-rich interactions, as SkillWeave combines teleoperation demonstrations for task-level coverage with kinesthetic demonstrations at contact-rich phases.

  2. The policy composition can reliably connect independently trained sub-task policies by introducing a bridging motion, which guides the robot from the predecessor’s realized terminal state toward the region of states represented by the successor’s demonstrated initial state distribution.

  3. The system can be robust to visual mismatch during kinesthetic learning because it utilizes an object-mask-conditioned diffusion policy, which uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting.

  4. The system mitigates distribution shift between sub-task policies through successoraware terminal steering, which selects actions to guide the system toward states supported by the successor’s demonstrated initial-state distribution, thereby achieving an average composition efficiency of 87%.

  5. The learned policy can be made more robust across different scenes and backgrounds because it uses object-mask conditioning instead of relying on online SAM 3 or image inpainting, which removes visible embodiments from kinesthetic demonstrations.

Sources

Related papers