Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors".
Rosa: Reinforcement learning methods typically learn new tasks from scratch, often disregarding prior knowledge that could accelerate the learning process.
Dev: First, who's behind it and why it matters.
Paper summary: Rosa: So, we're looking at this paper, "Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors," and it seems like the main idea is tackling the way reinforcement learning usually learns new things from scratch without any existing knowledge.
Dev: Right, Rosa, that's what drew me in; the abstract says RL methods often ignore prior knowledge, and this work proposes a method that models skill motions as having an unknown number of underlying features instead of sticking to a single Gaussian distribution for skills.
Taro: That makes sense from an autonomy standpoint because when things go wrong in complex tasks, you need a system that can adapt its underlying capabilities dynamically rather than just failing because it didn't learn the specific sequence it needed.
Rosa: Exactly, and what they claim is they use a Bayesian non-parametric model, specifically Dirichlet Process Mixtures enhanced with birth and merge heuristics, to pre-train this skill prior so that the system captures a diverse nature of skills.
Dev: That sounds like a flexible way to represent skills because it lets the model create new components or merge existing ones as it sees more data, which addresses the rigidity you mentioned earlier.
Taro: If they can dynamically capture an unknown number of features, that means the policy won't be locked into a fixed set of movements, which is crucial when dealing with unpredictable real-world interactions where the environment might misbehave.
Rosa: And they claim this pre-trained prior is then integrated into a hierarchical reinforcement learning framework to handle those extended long-horizon manipulation tasks.
Dev: The structure sounds promising for handling long sequences of actions, but I wonder about the computational cost; how does the loop rate hold up when you're feeding in these dynamic skill embeddings from that prior?
Taro: That brings up a point about robustness; if the system relies heavily on this learned skill prior, what happens if it encounters a completely novel situation that doesn't fit any of those seven identified base skills?
Rosa: The paper suggests they have achieved substantial improvements over state-of-the-art baselines in terms of average reward on long-horizon manipulation tasks, with some results stabilizing above three point seven and reaching up to four point zero in the Franka Kitchen environment <ref:2503.21975#pg0>.
Paper summary: Dev: A reward of that magnitude on those complex tasks is definitely something worth watching; it shows that this approach actually translates into better performance when dealing with extended sequences of actions, not just short demonstrations.
Taro: The fact that they successfully handle unseen subtasks during testing, like "open slider cabinet" and "open hinge cabinet," suggests a level of generalization we haven't seen before because the prior structure is so adaptable.
Rosa: That zero-shot skill adaptation capability really speaks to the power of modeling skills non-parametrically; it implies the prior isn't just memorizing movements but understanding the underlying mechanics.
Dev: From an engineering standpoint, that flexibility is great, but I'm curious about how they manage that exploration when using this KL divergence term instead of traditional entropy in their maximum entropy RL framework.
Taro: Replacing the standard policy entropy with a KL divergence term between the policy distribution and their pre-trained skill prior is an interesting way to guide exploration towards skills that are actually useful for the task at hand.
Rosa: It sounds like they're essentially telling the AI, "Don't just explore randomly; explore movements that align with what we already know about effective skills."
Dev: That alignment should improve exploration efficiency, but I need to know if this mechanism introduces any significant latency issues or failure modes when the system has to rapidly switch between those learned skill embeddings.
Taro: If the system is designed correctly, aligning the policy with a rich prior should help it navigate complex action sequences more intelligently than relying solely on trial and error for every single step of that long horizon.
Rosa: So, we're talking about a framework called "Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors" that uses these priors to guide hierarchical RL for tasks where learning from scratch is too slow.
Dev: It seems like the core thesis is using this flexible prior to replace fixed skill structures, which they achieve by modeling skills as having an unknown number of underlying features.
Paper summary: Taro: The implication here for autonomy is that we can build systems that don't just execute pre-programmed routines but can adapt their fundamental movement primitives based on the specific challenges they face during operation.
Rosa: If this framework works outside the lab, I wonder how long it could maintain that high performance when deployed in a messy, unstructured real-world setting compared to controlled simulations.
Dev: That's a big question for me; deployment success hinges on how well those learned skills generalize beyond the structured training environment where they were pre-trained.
Taro: The paper suggests they can bypass the need for reward annotations specific to every single task because the framework learns skills offline from unstructured action patterns, which opens up possibilities for real-world application across different scenarios.
Rosa: So, in simple terms, this paper presents a way to give robotic systems a flexible library of potential movements that they can draw from when tackling very long and complicated physical tasks.
Dev: It's about building an adaptive skill foundation rather than just training a single policy for the entire sequence of actions from start to finish.
Taro: The impact could be significant in areas where robots need to perform complex, multi-stage manipulation that isn't easily scripted beforehand, allowing for much more versatile physical interaction.
Rosa: The authors explicitly mention their limitation is that they are still working within a framework designed for specific long-horizon manipulation tasks, so generalizing this exact prior structure to vastly different domains is something they are looking toward.
Dev: That's fair; the current focus seems tightly bound to these types of sequential manipulation problems, which means we might need further work to see if this architecture scales easily into entirely different kinds of robotic control loops.
Taro: The paper's conclusion points toward combining this skill prior approach with foundation models for enhanced scalability in real-world long-horizon tasks, which suggests the next major step is integrating these priors into much larger, more general AI models.
Rosa: That sounds like a very exciting direction; linking this specific skill modeling to foundation model architectures could really unlock much broader applicability across different robotic domains.
Conclusion: Rosa: So, looking at the title of "Robotic Long-Horizon Manipulation with Bayesian Non-parametric Skill Priors," it really highlights how they're moving away from learning every single action from scratch and instead building a flexible library of movements first.
Dev: I agree, Rosa; that focus on skill priors suggests a system that learns fundamental abilities rather than just memorizing long sequences of actions for one specific task.
Taro: Exactly, and when you consider the authors who developed this work, they've clearly put a lot of thought into how to structure these priors so they can actually adapt when things get messy in the field.
Rosa: And those authors seem pretty confident because their results show that this framework handles unseen subtasks very well, which is huge for deployment outside of a perfectly controlled lab setting.
Dev: That generalization capability is something I'm curious about from an engineering standpoint; how long do you think this system can maintain that level of performance if it encounters unexpected physical friction or sensor noise in a real-world environment?
Taro: That's the million-dollar question, Dev; the paper itself points toward combining this skill prior approach with foundation models for better scalability, which suggests the next big step is making these skills robust enough for messy reality.
Rosa: It seems like this work is positioning us to build robots that can handle complex physical interactions without needing perfect pre-programming for every single scenario.
Dev: If we can get that loop rate and latency under control while leveraging these dynamic skill embeddings, it could seriously improve the reliability of autonomous systems performing intricate assembly or navigation tasks.
Taro: The implication is that autonomy in physical manipulation will become less about brute-force learning and more about intelligently selecting from a rich, learned set of underlying motor primitives.
Technical University of Munich
cs.RO, cs.AI
Submitted: 2025-03-27
Updated: 2026-10-03
Comments: Accepted by IROS 2026 (8 pages)
Project page: https://ghiara.github.io/HELIOS
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Reinforcement learning methods typically learn new tasks from scratch, often disregarding prior knowledge that could accelerate the learning process.
Key concepts
- Bayesian Non-parametric Skill Prior
- This is a flexible way to model skills where the number of underlying features is unknown beforehand. It uses a Dirichlet Process Mixture (DPM) model to dynamically create or merge skill components as new data arrives, allowing the system to capture an infinite range of potential behaviors rather than being limited by fixed skill sets.
- Hierarchical RL
- This structure organizes learning into two parts: an upstream module that learns high-level latent skill embeddings and a downstream component that uses these embeddings to generate a sequence of low-level actions. This allows the system to tackle long-horizon tasks by breaking them down into manageable skill steps.
- Skill Prior Pretraining
- This initial phase trains a model using an unstructured dataset (like video) to learn how skills are structured. It uses a Variational Autoencoder (VAE) and DPM to capture the non-parametric nature of skills, ensuring the learned prior is rich enough to represent diverse movements.
Terminology
Summary
Reinforcement learning methods typically learn new tasks from scratch, often disregarding prior knowledge that could accelerate the learning process. This work introduces HELIOS, a framework that integrates a Bayesian non-parametric skill prior with hierarchical RL to address long-horizon manipulation tasks, demonstrating superior performance and generalization compared to state-of-the-art baselines.
The gist
HELIOS is a framework that utilizes Bayesian non-parametric skill priors to guide downstream reinforcement learning policy learning for long-horizon manipulation tasks.
Framework Overview
The HELIOS framework is organized into two main phases: skill prior pretraining and RL for long-horizon manipulation. Phase I focuses on training a flexible, non-parametric skill representation model from an unstructured dataset, while Phase II deploys this pre-trained prior within a hierarchical RL framework to address complex, extended long-horizon tasks. The core idea is to model primitive skill motions as having non-parametric properties with an unknown number of underlying features,
which allows the system to capture a wider range of potential behaviors with a dynamic adaptive structure.
Phase I: Skill Prior Pretrain
This phase involves training a Bayesian non-parametric skill prior using a Dirichlet Process Mixture (DPM) model within the latent representation space. The process is iterative:
-
The VAE with GRU-based backbone models the non-parametric distributions of temporally extended skills based on an unstructured dataset.
-
The DPM is used to capture the
non-parametric nature of skills,
allowing for a dynamic structure wherenew components can be created or existing ones merged to improve the ELBO.
-
The total skill prior learning objective includes three components: reconstruction error, KL divergence between the skill posterior and Gaussian components, and KL divergence between the predicted skill prior and the true skill posterior (Eq. 7). This aims to
capture the non-parametric nature of skills, allowing the model to represent an unknown number of skill features that can adapt based on observed data.
Phase II: RL for Long Horizon Manipulation
The pre-trained skill prior is integrated into a hierarchical policy structure for downstream learning.
-
An upstream inference module, implemented with a Soft Actor-Critic (SAC) framework, learns latent skill embeddings based on observations and the pre-trained non-parametric skill prior.
-
The downstream component utilizes the
pre-trained skill decoder p(aiz,st) from Phase I,
which takes the latent embeddings and outputs a sequence of actions of length L, continuing until the next state update. -
To guide exploration in skill space, the traditional entropy term in the maximum entropy RL framework is replaced with a KL divergence term:
the policy entropy term is effectively equivalent to the negative KL divergence between the policy distribution and our pre-trained skill prior.
This encourages alignment with the skill prior distribution, improvingexploration efficiency and adaptability.
Experimental Validation
Experiments were conducted in the Franka Kitchen environment using a dataset from D4RL. The results demonstrated substantial improvements over baselines:
(Comparison)
HELIOS substantially outperforms all baseline models in terms of average reward on long-horizon manipulation tasks, quickly reaching a high average reward stabilizing above 3.7 (with maximal 4.0). HELIOS is shown to be superior to SPIRL because it leverages a more flexible Bayesian nonparametric skill prior modeled with DPM and MemoVB heuristics,
which allows it to dynamically capture an optimal set of skills that are more expressive and well-suited for complex, long-horizon tasks.
(Skill Representation)
The DPM-based prior effectively identifies and clusters the underlying features in the data, stabilizing around 6–8 components in the skill prior space. The final representation shows that our Bayesian non-parametric prior captures seven base skills: 'pick', 'place', 'pull', 'rotate', 'toggle', and 'explore' movement[s].
(Generalization)
The framework demonstrated zero-shot skill adaptation capability by successfully handling unseen subtasks, such as “open slider cabinet” and “open hinge cabinet,” highlighting the generalization capability of our framework
to apply prior knowledge to previously unencountered tasks.
Conclusions
HELIOS is a framework that integrates a Bayesian non-parametric skill prior with hierarchical RL to address long-horizon manipulation tasks, showing superior performance and generalization ability compared to state-of-the-art baselines. Future work aims to extend HELIOS by combining it with foundation models for enhanced scalability in real-world long-horizon tasks.
References
[1] K. Pertsch et al., “Accelerating reinforcement learning with learned skill priors,” in Conference on robot learning. PMLR, 2021, pp. 188–204.
[3] M. C. Hughes and E.
Improvements for AI systems
Here are specific improvements to existing AI systems based on the HELIOS framework, and what those improved systems can achieve:
-
The ability to learn and execute complex, multi-goal robotic manipulation tasks in long horizons (e.g., a sequence of 7 distinct subtasks like picking an object, placing it, toggling a switch) without requiring extensive task-specific training data from scratch.
-
Enhanced skill transfer efficiency across different manipulation environments or objects by leveraging a non-parametric representation of primitive skills rather than fixed models (like single Gaussians).
-
Improved exploration efficiency in long-horizon reinforcement learning by using the pre-trained Bayesian non-parametric skill prior to guide the policy's action selection, encouraging it to combine known fundamental actions in novel ways.
-
Increased zero-shot adaptation capability: The improved system can successfully learn and execute entirely new subtasks (like
open slider cabinet
orturn on top burner
) that were never present in the initial training dataset, demonstrating true generalization of primitive skill knowledge. -
More interpretable skill representation: The framework provides a structured latent space (via t-SNE projections) where distinct primitive skills are clearly clustered, allowing researchers to understand exactly which underlying motions (e.g.,
pick,
rotate
) the agent has learned and how they are being recombined. -
Robust policy learning under sparse rewards: The system can achieve high performance even when trained on sparse, episodic rewards characteristic of long-horizon tasks, where traditional methods often fail to converge or explore effectively.
Abstract
Long-horizon manipulation under sparse rewards remains challenging for reinforcement learning due to delayed feedback and inefficient exploration. Existing skill-based approaches often assume a fixed parametric prior (e.g., a single Gaussian), limiting their ability to capture diverse and multi-modal skill structures required for complex tasks. We propose a Bayesian non-parametric skill prior that models temporally extended skills in a structured latent space using a Dirichlet Process Mixture, enabling adaptive skill discovery without predefining the number of components. Integrated into a hierarchical RL framework, the learned prior guides high-level skill selection while a pretrained decoder generates temporally abstracted actions, improving exploration efficiency in sparse-reward settings. Experiments on Franka Kitchen, LIBERO-Long, Meta-World, and a real robot demonstrate consistent gains in long-horizon manipulation, achieving over 0.8 success rate within 1.5M steps, whereas SAC fails to converge even after 5M steps (< 0.1). Compared to a single-Gaussian prior baseline, our model yields an average improvement of 21.8%.
Sources
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning
- Stick-Breaking Variational Autoencoders
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving