RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
summary
The gist
RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills
In short
RoboHarness is a framework that allows a central coding agent to plan long-horizon tasks using diverse, independent robot policies. It achieves this by wrapping each policy as an 'agentic skill' and employing memory skills to intelligently route between them. This enables stable handoffs between different robot capabilities without needing joint retraining or shared representations, improving planning for complex, varied tasks.
Key concepts
- Agentic Skill Modules
- These are independently developed or hand-designed robot policies that RoboHarness wraps. They act as reusable skills for the main planner. Each module has a 'policy card' detailing its specific capabilities and constraints, allowing the system to treat different robots or skills as interchangeable tools.
- Understanding Skills
- These skills interpret raw task inputs to assess how well current states match policy capabilities. They perform five checks: assessing temporal stability of poses, comparing visual observations to training data, matching task instructions to policy descriptions, checking state compatibility with the next policy's region, and evaluating input quality.
- Memory Bridge
- This mechanism manages execution history through a memory bank. It retrieves relevant past experiences based on both text and visual similarity. It then constructs a local support region around retrieved memories and scores candidate states to determine the best target state for the next policy transition.
- Self-Evolution Skills
- These skills allow RoboHarness to adapt online using execution evidence. They include modifying the overall orchestration structure, tuning hyperparameters like retrieval settings, and updating each policy's capability metadata based on real-world successes and failures.
Terminology used across episodes
This episode discusses
- RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning · Paper Radio
- From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- pi 0: A Vision-Language-Action Flow Model for General Robot Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Tree-Planner: Efficient Close-loop Task Planning with Large Language Models
- One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model · Paper Radio
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
- pi* 0.6: a VLA That Learns From Experience
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization
- Causal World Modeling for Robot Control
- RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
- Learning to Compose Skills
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
The paper
RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning · Read on arXiv
Jinbang Huang, Yuanzhao Hu, *Zhiyuan Li*, *Ran Qi*, Yixin Xiao, Zhanguang Zhang, Mark Coates, Tongtong Cao, Yingxue Zhang
Huawei Noah’s Ark Lab
Long-horizon robotic tasks require a breadth of capabilities beyond what any single existing robot control policy can reliably provide. Combining heterogeneous policies with complementary strengths offers a promising solution, but introduces two key challenges: uncertain capability boundaries and distribution mismatches during policy handoffs. These challenges remain largely unaddressed by existing planning methods, which typically assume homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed heterogeneous policies, including vision-language-action models (VLAs), world-action models (WAMs), reinforcement learning (RL) policies, and task and motion planners (TAMP), as reusable agentic skills. RoboHarness integrates understanding, memory, and evolution skills to reason about policy capabilities and support capability-aware task decomposition and policy routing. To mitigate distribution mismatches during policy handoffs, we introduce Memory Bridge, a plug-in policy-chaining mechanism that enables reliable transitions between heterogeneous policies without joint retraining. Extensive experiments across five public benchmarks, 500 customized tasks across 10 classes, and 135 real-robot trials demonstrate substantial gains in long-horizon and memory-dependent tasks, as well as robustness to out-of-distribution conditions.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning".
Rosa: RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills and fixed applicability.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So, looking at "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning," it seems like the core idea is moving away from static skill spaces where everything is pre-defined, which is what most existing long-horizon planning approaches rely on.
Dev: The authors are Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Zhanguang Zhang, Mark Coates and Tongtong Cao. I’m looking at their background; they seem to have a mix of expertise in robotics and foundation models.
Taro: I see the authors have experience that spans different areas of autonomy research; this suggests the work might bridge the gap between high-level reasoning and low-level control effectively.
Rosa: That’s true, Taro, and what’s important is that they are proposing a unified framework rather than just tweaking one existing planning algorithm to handle these heterogeneous systems.
Dev: The authors argue that because policies differ in architecture and history, we can't assume their capability boundaries are clear or fixed for every task context.
Taro: That uncertainty about when one policy stops being reliable is a key challenge in real-world autonomy, so addressing that directly seems like the right direction here.
Rosa: And they introduce this concept of using memory to manage these handoffs, which is the mechanism they focus on to make the orchestration possible across different skill types.
Dev: It shifts the focus from just planning a sequence of actions to reasoning about which specific policy should take over at any given moment based on learned context.
The paper's summary: Rosa: In terms of what RoboHarness actually does, it proposes an agentic orchestration system where a coding agent acts as the high-level planner and router to manage these various robot policies, treating each policy as a distinct skill module.
Dev: It builds on this by wrapping those modules with three auxiliary skills—Understanding, Memory, and Self-Evolution—to provide the necessary information flow for this capability-aware planning.
Taro: So the Understanding skills are responsible for interpreting what's happening in the environment and figuring out how that relates to which policy might be capable of handling it next.
Rosa: Precisely, and those understanding skills include things like assessing uncertainty in object poses and checking if the current state is within the known operating region of a candidate policy.
Dev: Then there's the Memory component, which uses a Memory Bridge to store execution histories in a structured way, allowing it to retrieve relevant past experiences when deciding on a transition between policies.
Taro: The memory bridge sounds like it’s crucial for maintaining spatial consistency during these handoffs, ensuring that even if we switch policies, the robot doesn't suddenly end up in an impossible configuration.
Rosa: And finally, the Self-Evolution skills let the system adapt online by refining its routing strategies and tuning its own parameters based on execution evidence it gathers.
The paper's improvements: Dev: The paper suggests several key improvements over previous methods, primarily focusing on how to handle distribution mismatch and capability boundaries between these different policies during planning.
Rosa: One major improvement is the move towards capability-aware decomposition and routing, meaning the planner doesn't just pick a skill; it reasons about which policy is best suited for that specific situation right now.
Taro: That dynamic routing should solve a lot of issues with static task decomposition where you have to guess the right skill before execution even starts.
Dev: Another improvement centers on creating a stable inter-policy handoff mechanism, specifically through the Memory Bridge, which uses both semantic and visual similarity to preserve spatial continuity.
Rosa: That bridge is what allows them to maintain progress even when transitioning between policies whose input and output distributions don't naturally align.
Taro: It sounds like this structure helps solve the problem where a robot might get stuck because the next intended action simply isn't in the policy's learned distribution.
Dev: The Self-Evolution skills offer an improvement by letting the system improve its own orchestration rules and parameter settings through online execution feedback, making it more robust over time.
Conclusion: Rosa: To wrap up our discussion on "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning," the main implication is that we can build systems capable of handling complex, multi-step tasks that require diverse, specialized abilities.
Dev: It moves us past the limitations of assuming all skills are homogeneous by introducing a dynamic way to match the task needs to the specific capabilities available in a collection of different robot control systems.
Taro: For me, it means we can expect autonomy to get much more resilient when facing unexpected environmental changes because it has these mechanisms for on-the-fly adaptation and recovery.
Rosa: It really suggests that instead of trying to build one massive, monolithic policy, we can construct a system from smaller, independently developed components that work together intelligently.
Dev: That intelligence is driven by the memory and understanding layers that manage the flow of information between those distinct components so they don't just execute in isolation.
Taro: I think the real impact is enabling more generalist agents that can tackle novel problems by intelligently combining existing tools rather than requiring a completely new, monolithic architecture for every single application.
Rosa: That’s a strong summary of how RoboHarness aims to improve long-horizon planning by managing the complexity of heterogeneous systems.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications