Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning

arXiv:2606.15685 · cs.RO, cs.CV · Submitted 2026-06-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Learning New Tasks via Reusable Skills".

Dev: Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Welcome back to the show. We're diving into some interesting work today on embodied continual learning. We've got a paper called "Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning." Rosa, what are your initial thoughts on this title and who the authors are?

Dev: I’m curious about the mechanics behind how they tackle that challenge of continual learning without catastrophic forgetting, Rosa. Who were the main researchers on this project?

Taro: I'm looking forward to hearing about how they handle those real-world deployment scenarios, Rosa. Does this work hold up outside of a controlled lab environment for any significant duration?

Rosa: Well, this paper proposes a framework called Skill-Compositional Experts for Embodied Continual Learning. The authors are Shuaike Zhang, Shaokun Wang, Haoyu Tang, Jianlong Wu, Liqiang Nie. They’re coming from institutions like Harbin Institute of Technology and Shandong University.

Dev: That's a solid team of researchers behind it; I wonder if their background in control engineering gives them an edge here regarding the closed-loop aspect. Rosa, can you give us a simpler way to understand what this paper is trying to achieve?

Rosa: Essentially, they are trying to solve a big problem where robots have to learn new manipulation tasks one after another while keeping everything they learned before still working properly under active control. They argue that current methods cause feature drift because the robot's understanding of the world keeps shifting toward the newest task. This paper introduces a way to organize skills so the robot doesn't lose old knowledge when it learns something new.

Dev: So, instead of treating every task as completely separate, they are suggesting a system where skills are reusable and can be combined to make new things. That sounds like it could simplify how we think about complex robotic behavior over time.

Taro: From an autonomy researcher's viewpoint, I'm really interested in the composition part. How does this framework handle situations where the environment throws something completely unexpected at the robot during a task sequence?

Rosa: The paper explains that they build a skill base using something called Compositional Skill Grounding, or CSG. This component breaks down demonstrations into reusable skills and organizes them into that base so knowledge can be reused across different tasks.

Dev: That skill base sounds like a library of building blocks for actions, which is much better than just retraining the whole model every time we want to do something different. Rosa, what's the core mechanism they use to actually implement this skill reuse?

Rosa: The core mechanism involves Dual Execution-and-Transition Experts, or DETE. This system augments the main VLA action decoder with two distinct branches: one that handles how to execute a specific reusable skill, and another that models the patterns for switching between those skills when things change.

Taro: That sounds like it directly addresses the problem of feature drift by separating what's happening during execution from what happens when the robot is deciding which skill to switch to next. How does that separation actually work in practice?

Title and authors: Dev: It uses a sample-level skill distribution derived from the conditional token, which they call p = softmax(MLP(c)), and then use that to select an execution expert, = arg max p k. This means when the system is focusing on one skill, it strictly uses the expert for that skill.

Rosa: And around the boundaries between those skills, they use a transition-aware routing feature h = MLP(c) + Wpp to figure out how much to rely on those cross-skill patterns. This dynamic balancing is managed by an Adaptive Fusion Module that decides the final action y by mixing the execution and transition outputs.

Taro: So, if I'm testing a robot in a sequence, and it needs to switch from grasping something small to moving something large, this mechanism should allow it to handle that switch without completely breaking its control loop. What happens when the skill distribution p is very spread out across many skills?

Dev: If p is spread out, the transition expert branch becomes more influential because the system can't confidently pick just one execution expert. The Adaptive Fusion Module adjusts its weight alpha, shifting emphasis from pure execution to transition modeling. This keeps things flexible when the environment is ambiguous.

Rosa: The paper shows that in their experiments on LIBERO, they achieved a success rate of eighty-two point five percent for the overall performance compared to about seventy-three point seven percent for the existing T-MoE method, and they even managed a lower forgetting rate of four point three percent after closed-loop control compared to fifteen point eight percent.

Taro: That difference in forgetting rates is quite significant when you consider the long sequence of tasks they tested in that benchmark. It suggests that this skill-compositional approach really helps maintain old task knowledge during extended operation.

Dev: The performance metrics show a clear gain in robustness, but I'm still looking at the latency implications of running both expert branches simultaneously on real hardware. They mention that the design confines execution updates to the expert associated with the current skill, which should help keep those local feature drifts manageable for our control loop requirements.

Rosa: The paper also confirms that both parts are necessary; removing CSG significantly degrades performance, and taking away either the Execution Expert Branch or the Transition Expert Branch lowers metrics like AUC and Final SR, showing they complement each other.

Taro: That separation between execution fidelity and transition modeling seems key to achieving that stability-plasticity trade-off they mentioned. It means we get a better balance between sticking to what we know and being flexible enough to adapt when things get messy.

Dev: If the system is relying on skill composition, does this mean the robot can learn truly novel behaviors that it hasn't seen before by just putting known skills in a new order? That’s a big question for practical deployment.

Title and authors: Rosa: Exactly, because they demonstrate that new tasks can be composed from reusable skills in the skill base, showing they can execute sequences like "Bowl in Top Drawer" by combining existing skills. It suggests a pathway to learning complex, multi-step goals without needing a completely fresh dataset for every single novel task.

Taro: If this framework proves effective under closed-loop control for long periods, the implications could be huge for deploying mobile robots in unstructured environments where tasks are constantly changing and we can't just stop and retrain.

Dev: I still need to see how stable the training process is when you introduce a brand new skill into that skill base during continuous operation. That dynamic updating of B needs to be very robust for real-time systems, Rosa.

Rosa: The paper suggests they are focused on making that grounding process efficient through their use of VisionLanguage Models to ground observations and language instructions with respect to the current skill base. They are trying to make sure the skill base grows intelligently and reliably.

Taro: So, while the theoretical structure is powerful, I'm keen to see how well it handles unexpected physical interactions that don't fit neatly into pre-defined skills. What happens when things misbehave in a way that no existing skill covers?

Dev: The paper addresses this by proposing a mechanism where if no existing skill matches the observation, the system adds it to the base B as a new reusable skill, which is how it handles novelty during operation. That's an explicit way to incorporate unforeseen skills.

Rosa: It seems like they are providing a structured way for robots to evolve their capabilities incrementally rather than suffering from complete knowledge loss when encountering something entirely new. This structure is what sets Skill-Compositional Experts apart in the field of embodied continual learning.

Taro: I think the ability to compose sequences and adapt dynamically through DETE gives us a much more resilient agent for complex, long-horizon tasks that are common in real-world manipulation scenarios.

Dev: We need to keep watching their work on how they manage the update frequency of that skill base B during sustained operation. The performance gains look impressive, but stability under high operational load is where the engineering challenge will lie for us.

Rosa: Well, that’s what we’ll be looking at in the next part. We've covered a lot about how this paper addresses feature drift and composition. Now we're going to wrap up and get ready for our next topic.

Taro: I think the resilience they show in maintaining old task knowledge while learning new ones through skill composition is a really important direction for future embodied AI research.

Dev: Agreed, the stability trade-off they found seems very practical for real-world applications where reliability is paramount.

Rosa: And that's our time for this discussion on "Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning." We’ll be right back after the break with a look at some work on power systems.

The paper's summary: Rosa: So, we've looked at how they structure skill learning through reusable skills and then we're going to walk through what their summary actually says about this Skill-Compositional Experts framework and where this stuff could take us next.

Dev: I’m ready for the summary; I want to make sure I understand the core mechanism they propose for handling that feature drift we talked about earlier.

Taro: I'm eager to hear how their composition idea translates into something that can actually handle unpredictable situations in the physical world, not just simulated ones.

Rosa: Alright, so this paper essentially lays out Skill-Compositional Experts as a way to stop robots from forgetting what they learned while learning new things by making sure old skills are organized into a reusable library and then allowing the robot to build entirely new tasks by combining those existing blocks.

Dev: That sounds like they're moving away from learning every single task in isolation, which is exactly what we need for sustained operation under closed-loop control. They suggest that instead of retraining the whole model when a new manipulation goal comes up, the system just grounds that new instruction into the existing skill base and uses those skills to build something new.

Taro: I like that idea of composition; it’s not just about incremental learning anymore, it’s about building complex behaviors by stringing together what we already know. That implies a robot could learn a whole sequence of actions just by combining "grasp," "move," and "open" skills in a novel order.

Rosa: Precisely, Taro; they show that this composition is powerful for new tasks, and the framework uses Dual Execution-and-Transition Experts to manage the execution fidelity versus the necessary skill switching. This means when things get complicated, it can execute a known skill perfectly while simultaneously figuring out how to transition smoothly if it needs to move into a different part of its task library.

Dev: The detail about those two expert branches is interesting because it’s like having two specialized controllers running in parallel; one focused purely on doing the current action and the other focused on managing the handoff between actions. That structure should help keep our loop rates stable, even when the skill distribution p starts shifting rapidly.

Taro: If that transition modeling is robust, it means we could see agents operating in really messy, dynamic environments where tasks are constantly changing their requirements on the fly, and they wouldn't just crash or revert to old behavior because they lost track of how to switch strategies.

Rosa: That resilience is what makes this framework so exciting; it suggests a path toward robots that can truly evolve their capabilities over time rather than just being static learners for one specific set of goals. The results on the LIBERO benchmarks show they keep performance high even as the complexity increases, which points toward solid long-term viability.

Dev: I'm still focusing on the closed-loop aspect; if we run this system in a real factory setting, how much computational overhead is added by running both those expert branches simultaneously for every single decision cycle? We need to keep that latency low for reliable control.

Taro: That’s a valid concern, Dev; the efficiency of that Adaptive Fusion Module is crucial. If they can dynamically shift the weights so that execution dominates when things are stable and transition modeling takes over when things are ambiguous, it keeps the computational load manageable without sacrificing safety or coherence.

Rosa: It sounds like we have a solid foundation here—a way to structure skill knowledge for continuous learning through composition and expert-level decision-making. Next up, we'll look at how this impacts the broader field of autonomous systems and where researchers see the next steps for deploying these skills in real-world scenarios.

The paper's improvements: Tom: So, we've seen how they structure skill learning through reusable skills and now we’re going to talk about the specific improvements they propose for this Skill-Compositional Experts framework and what that actually means for how these robots operate in the field.

Dev: I'm interested in the practical fixes; what exactly are they suggesting to make the control loop more robust against those feature drifts we discussed?

Taro: I'm curious about the mechanisms they suggest for handling those unexpected situations, especially when a robot encounters something it hasn't seen before during a skill sequence.

Rosa: To recap, the paper outlines improvements that focus on explicitly structuring knowledge into reusable skills and using Dual Execution-and-Transition Experts to manage how the AI handles execution versus skill switching dynamically.

Dev: The main improvement is in how they handle the action decoder; they’re augmenting it with those two expert branches so that execution updates stay tightly focused on the current skill, which should really help keep our latency predictable during high-speed maneuvers.

Taro: I see that separation as a huge safety feature; if one branch handles the precise motion of a known skill and the other manages the transition patterns, it gives us a clear way to monitor for errors in either domain. That predictability is what autonomy researchers look for when dealing with real-time uncertainty.

Rosa: Exactly, Taro; that structure allows the system to maintain high fidelity during execution while remaining flexible enough to pivot quickly if environmental features start drifting away from the expected patterns of that skill. This directly tackles the core issue of feature drift in a closed-loop setting.

Dev: I’m looking at the transition modeling part too; they propose making that cross-skill pattern calculation more aware of the current skill distribution p, which means when p is concentrated, it emphasizes execution over switching, and when p is spread out, it ramps up the cross-skill logic. That adaptive weighting sounds like a smart way to manage computational load while maintaining necessary flexibility.

Taro: If that adaptive weighting works as intended, we could have agents that are incredibly stable during routine tasks but possess the necessary agility to switch strategies instantly when faced with novel or highly unpredictable physical interactions. It moves beyond just learning sequences and toward true adaptive problem-solving.

Rosa: That’s the big picture; it points toward a future where embodied systems aren't just following pre-programmed paths but are actively composing solutions on the fly as they explore new situations in the real world. This is what makes me optimistic about its potential outside of a clean lab environment, though we still need to see how long that skill base B can grow before it becomes computationally unmanageable.

Dev: That growth management is the sticking point for me; if the skill base keeps expanding too quickly with every new object or interaction, the system’s decision-making engine will slow down, and that defeats our purpose for a fast control loop. We need to ensure CSG is efficient enough to keep adding skills without creating a bottleneck.

Taro: I agree with Dev; the scaling of that grounding component is critical; if it takes too long to ground and add a new skill, the system loses its ability to react quickly in dynamic scenarios. The implication here is that for this technology to be useful in real-world deployment, the skill discovery process itself has to be highly efficient and fast.

Rosa: So, we have a framework that structures knowledge compositionally and uses expert branching for execution and transition management, which promises much better stability than current methods. We’ve seen how it can handle long sequences without forgetting old stuff, but we need to keep pushing on the efficiency of skill discovery so these robots can operate reliably for extended periods in messy environments.

Conclusion: Rosa: So, to wrap things up, we’ve seen how the Skill-Compositional Experts framework tackles feature drift in embodied continual learning by using compositional skill grounding and dual execution-and-transition experts for robust task acquisition.

Dev: It really lays out a clear architecture where execution fidelity and skill switching are modeled separately but fused adaptively, which is exactly what we need to keep the control loop tight during continuous operation.

Taro: From an autonomy standpoint, this framework suggests that robots won't just get stuck when the world throws something unexpected at them because they can compose new actions from known parts of their skill library.

Rosa: The implication is that we could see mobile robots operating in complex, long-horizon environments for extended periods without the degradation of performance we’ve seen in other continual learning methods.

Dev: I'm still thinking about the operational longevity; if the skill base keeps growing too fast due to new observations, we have to be extremely careful about how that scaling impacts our real-time constraints and failure modes.

Taro: That's a valid point, Dev; the efficiency of that skill discovery mechanism is what separates theoretical potential from practical deployment in a constantly changing physical world.

Rosa: Exactly, Taro; it’s not just about learning one task well, but about building an adaptable system that can handle infinite novel tasks by composing existing ones intelligently. This is what makes the Skill-Compositional Experts framework so compelling for field robotics.

Dev: I think the separation of execution and transition modeling is the most important technical takeaway for me; it gives us a way to debug *why* a failure happened, whether it was a bad skill implementation or an unstable transition.

Taro: It moves us closer to building agents that can truly handle unforeseen physical interactions by having that structured approach ready when things go sideways.

Rosa: We’ve got some great insights from this discussion on Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning. Next up, we’ll be looking at how these principles apply to other areas of AI, specifically in the realm of robust control systems and power management.

Dev: I'm ready for that shift; I want to see how this skill composition idea might translate into better stability guarantees when we look at those control papers.

Taro: I'm looking forward to hearing how this framework impacts the broader autonomy landscape in the next segment.

Shuaike Zhang, Shaokun Wang, Haoyu Tang, Jianlong Wu, Liqiang Nie

Harbin Institute of Technology, Shenzhen · Shandong University · Shenzhen Loop Area Institute

cs.RO, cs.CV

Submitted: 2026-06-14

Updated: 2026-09-28

Comments: 12 pages, 4 figures, 3 tables

Project page: https://eqcy.github.io/sce

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 80/100

The gist: Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control, but existing methods

Key concepts

Compositional Skill Grounding (CSG)
This component breaks down task demonstrations into reusable skills. It first decomposes temporal segments based on gripper state changes or motion trends. Then, it uses a Vision-Language Model to ground these observations and instructions against the existing skill base, proposing new skills or selecting existing ones for reuse.
Dual Execution-and-Transition Experts (DETE)
DETE augments the VLA decoder with two expert branches. The Execution Expert models skill-specific actions for executing a chosen skill, while the Transition Expert models patterns for switching between different skills. An Adaptive Fusion Module dynamically combines these outputs to ensure stable execution and flexible handling of skill changes.
Skill Composition
The framework treats every new task as a combination of previously learned, reusable skills. This allows the system to leverage existing knowledge rather than learning entirely new representations from scratch for every task, which is key to mitigating feature drift in embodied continual learning.

Terminology

Summary

Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control, but existing methods suffer from severe catastrophic forgetting due to feature drift.

The gist: SCE is a Skill-Compositional Experts framework for ECL that builds a skill base via Compositional Skill Grounding (CSG) and enables new task learning through skill composition using Dual Execution-and-Transition Experts (DETE).

Problem Definition

In embodied continual learning, a VLA model must sequentially acquire new manipulation tasks while retaining previously learned ones under closed-loop control. This setting is more challenging than conventional continual learning because feature representations of Vision-Language-Action (VLA) models progressively drift toward new tasks, gradually deviating from those associated with previously learned ones, and this accumulated feature drift propagates through sequential decision-making and eventually degrades the execution of previously learned tasks. Existing methods primarily mitigate catastrophic forgetting at the task or skill level, treating each task as an isolated learning unit. However, this approach is suboptimal for embodied manipulation because it limits the exploitation of shared patterns across tasks, which are essential for generalizing across diverse tasks.

Proposed Framework: SCE

The proposed framework is Skill-Compositional Experts (SCE), which introduces a compositional view of ECL where each task can be represented as a composition of reusable skills. SCE consists of two main components:

  1. Compositional Skill Grounding (CSG): This component decomposes task demonstrations into reusable skills and organizes them into the skill base, enabling reuse of skill knowledge across tasks. CSG proceeds in two phases: Phase 1 involves State-Aware Task Decomposition, where temporal segments are defined based on rules concerning gripper state changes or end-effector motion trends. Phase 2 uses a VisionLanguage Model (VLM) to ground the segment observation O(m) and the language instruction l with reference to the current skill base B, and proposes a candidate skill sˆ(m). If no existing skill matches, it is added to B as a new reusable skill; otherwise, the matched skill is selected.

  2. Dual Execution-and-Transition Experts (DETE): DETE augments the VLA action decoder with two complementary expert branches. The Execution Expert Branch models skill-specific action patterns for the execution of reusable skills, while the Transition Expert Branch models cross-skill transition patterns for skill switching during task execution. DETE employs an Adaptive Fusion Module to dynamically combine these outputs, ensuring stable skill execution while flexibly handling skill transitions, thereby reducing feature drift during ECL.

DETE Details and Skill Composition

DETE is integrated parallel to the pretrained action decoder. It operates by deriving a sample-level skill distribution p from the conditional token c using the formula: p = softmax(MLP(c)), Lskill = CE(p, s). This distribution p characterizes each input in terms of reusable skills. The Execution Expert Branch uses this distribution to select the most relevant skill: ˆk = arg max pk, ∆y exe = Ekˆ(y), where Ek is associated with the k-th dimension of the skill space V. This design confines execution updates to the expert associated with the currently involved skill, which reduces interference with unrelated skills and mitigates feature drift. The Transition Expert Branch models cross-skill dependencies by calculating a transition-aware routing feature h: h = MLP(c) + Wpp, which is then used to compute mixture weights u over N experts. Finally, an Adaptive Fusion Module computes the output action ∆y by adaptively balancing the branches: ∆y = α∆y exe + (1 − α)∆y tr. This mechanism ensures that when p is concentrated on a single skill, DETE emphasizes execution; around skill boundaries, variations in p modify h and shift the fusion toward cross-skill transitions.

Experimental Validation

Experiments on LIBERO benchmarks and real-world manipulation tasks demonstrate SCE's effectiveness. In simulation, SCE achieves lower forgetting while maintaining stronger overall performance, as shown in Figure 1(c). In real-world experiments, SCE shows superior results compared to baselines like Task-MoE. Feature drift analysis confirms that SCE mitigates old-task feature drift, especially at the action-feature levels, and maintains a smaller simulator-state gap from the reference behavior under closed-loop control. Ablation studies confirm that both CSG and DETE are crucial: removing CSG degrades performance significantly, and removing either expert branch (TEB or EEB) leads to lower AUC and Final SR, suggesting that skill-specific execution modeling and cross-skill transition modeling are complementary. The results indicate a stronger stability-plasticity trade-off for skill-compositional ECL.

Contributions

The primary contributions are:

Improvements for AI systems

Here are the specific improvements for AI systems based on the proposed SCE framework, and what these improved systems can achieve:


The proposed Skill-Compositional Experts (SCE) framework fundamentally addresses catastrophic forgetting in Embodied Continual Learning (ECL) by explicitly structuring skill knowledge, transforming task learning from isolated learning units into a compositional process.

Here are the specific improvements and capabilities:

  1. A system that can perform continuous, multi-task manipulation without forgetting old skills by leveraging a reusable skill base.

  2. The ability to learn entirely new manipulation tasks by composing existing learned skills together in novel sequences (skill composition).

  3. Robust performance under closed-loop control, where the robot must adapt its actions based on real-time environmental feedback, without significant degradation of previously mastered behaviors (mitigating feature drift).

  4. A mechanism for skill switching that is coherent and predictable, allowing the agent to transition smoothly between different manipulation strategies required for sequential tasks.

Specific technical improvements enabled by SCE:

  1. The system will utilize a dedicated skill decomposition module (CSG) to automatically parse raw task demonstrations into a library of reusable, primitive skills grounded in visual-language tokens and state dynamics.

  2. The action decoder will be augmented with the Dual Execution-and-Transition Experts (DETE).

  3. The system will maintain an Execution Expert Branch dedicated to executing the currently relevant skill, ensuring high fidelity for that specific motion pattern, thereby localizing feature drift and preserving skill integrity.

  4. The system will utilize a Transition Expert Branch that models cross-skill dependencies and transition patterns between skills, allowing it to generate coherent action sequences during task switching rather than erratic behavior.

  5. The Adaptive Fusion Module within DETE will dynamically weigh the outputs of the execution and transition branches based on the current skill distribution, ensuring that when executing a known skill, execution dominates, and when transitioning between skills, cross-skill modeling takes precedence.

What this improved AI system can do:

The resulting SCE-based robot/agent can perform complex manipulation in dynamic environments with high reliability across an infinite stream of novel tasks. Specifically:

  1. It can learn a new sequence of actions (e.g., Pick up object A, move it to location B, and then open container C) by seamlessly combining previously learned skills (grasping, moving, opening).

  2. When performing the combined task sequence under real-time control (closed-loop), it will execute the necessary skill components correctly, even if environmental features drift due to new objects or states, because the system knows exactly which skill branch to rely on for execution and which branch to use for smooth transitions.

  3. It can maintain high success rates (Final SR) across a long sequence of tasks (up to 10 stages in LIBERO benchmarks) without experiencing catastrophic forgetting of earlier skills, demonstrating superior long-term adaptability compared to current methods that treat every task as isolated.

Abstract

Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control. In ECL, feature drift can propagate through sequential decision-making under closed-loop control, turning representation changes into compounding behavioral deviations on previously learned tasks. A key challenge in ECL lies in structured skill reuse across continually evolving tasks, since existing methods primarily focus on skill learning without explicitly organizing them for coherent task execution. To address this issue, we propose SCE, a Skill-Compositional Experts framework for ECL. SCE builds a skill base via Compositional Skill Grounding (CSG), which decomposes task demonstrations into reusable skills. Based on this, Dual Execution-and-Transition Experts (DETE) enable new task learning through skill composition, where one branch ensures skill execution and the other supports transitions between skills for coherent behavior. Experiments on LIBERO benchmarks and real-world manipulation tasks show that SCE improves retention and overall task performance. Further feature drift analyses and ablation studies verify the effectiveness of our method. Project website: https://eqcy.github.io/sce/.

Sources

Related papers