RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning".
Rosa: RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills and fixed applicability.
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: So, looking at "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning," it seems like the core idea is moving away from static skill spaces where everything is pre-defined, which is what most existing long-horizon planning approaches rely on.
Dev: The authors are Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Zhanguang Zhang, Mark Coates and Tongtong Cao. I’m looking at their background; they seem to have a mix of expertise in robotics and foundation models.
Taro: I see the authors have experience that spans different areas of autonomy research; this suggests the work might bridge the gap between high-level reasoning and low-level control effectively.
Rosa: That’s true, Taro, and what’s important is that they are proposing a unified framework rather than just tweaking one existing planning algorithm to handle these heterogeneous systems.
Dev: The authors argue that because policies differ in architecture and history, we can't assume their capability boundaries are clear or fixed for every task context.
Taro: That uncertainty about when one policy stops being reliable is a key challenge in real-world autonomy, so addressing that directly seems like the right direction here.
Rosa: And they introduce this concept of using memory to manage these handoffs, which is the mechanism they focus on to make the orchestration possible across different skill types.
Dev: It shifts the focus from just planning a sequence of actions to reasoning about which specific policy should take over at any given moment based on learned context.
The paper's summary: Rosa: In terms of what RoboHarness actually does, it proposes an agentic orchestration system where a coding agent acts as the high-level planner and router to manage these various robot policies, treating each policy as a distinct skill module.
Dev: It builds on this by wrapping those modules with three auxiliary skills—Understanding, Memory, and Self-Evolution—to provide the necessary information flow for this capability-aware planning.
Taro: So the Understanding skills are responsible for interpreting what's happening in the environment and figuring out how that relates to which policy might be capable of handling it next.
Rosa: Precisely, and those understanding skills include things like assessing uncertainty in object poses and checking if the current state is within the known operating region of a candidate policy.
Dev: Then there's the Memory component, which uses a Memory Bridge to store execution histories in a structured way, allowing it to retrieve relevant past experiences when deciding on a transition between policies.
Taro: The memory bridge sounds like it’s crucial for maintaining spatial consistency during these handoffs, ensuring that even if we switch policies, the robot doesn't suddenly end up in an impossible configuration.
Rosa: And finally, the Self-Evolution skills let the system adapt online by refining its routing strategies and tuning its own parameters based on execution evidence it gathers.
The paper's improvements: Dev: The paper suggests several key improvements over previous methods, primarily focusing on how to handle distribution mismatch and capability boundaries between these different policies during planning.
Rosa: One major improvement is the move towards capability-aware decomposition and routing, meaning the planner doesn't just pick a skill; it reasons about which policy is best suited for that specific situation right now.
Taro: That dynamic routing should solve a lot of issues with static task decomposition where you have to guess the right skill before execution even starts.
Dev: Another improvement centers on creating a stable inter-policy handoff mechanism, specifically through the Memory Bridge, which uses both semantic and visual similarity to preserve spatial continuity.
Rosa: That bridge is what allows them to maintain progress even when transitioning between policies whose input and output distributions don't naturally align.
Taro: It sounds like this structure helps solve the problem where a robot might get stuck because the next intended action simply isn't in the policy's learned distribution.
Dev: The Self-Evolution skills offer an improvement by letting the system improve its own orchestration rules and parameter settings through online execution feedback, making it more robust over time.
Conclusion: Rosa: To wrap up our discussion on "RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning," the main implication is that we can build systems capable of handling complex, multi-step tasks that require diverse, specialized abilities.
Dev: It moves us past the limitations of assuming all skills are homogeneous by introducing a dynamic way to match the task needs to the specific capabilities available in a collection of different robot control systems.
Taro: For me, it means we can expect autonomy to get much more resilient when facing unexpected environmental changes because it has these mechanisms for on-the-fly adaptation and recovery.
Rosa: It really suggests that instead of trying to build one massive, monolithic policy, we can construct a system from smaller, independently developed components that work together intelligently.
Dev: That intelligence is driven by the memory and understanding layers that manage the flow of information between those distinct components so they don't just execute in isolation.
Taro: I think the real impact is enabling more generalist agents that can tackle novel problems by intelligently combining existing tools rather than requiring a completely new, monolithic architecture for every single application.
Rosa: That’s a strong summary of how RoboHarness aims to improve long-horizon planning by managing the complexity of heterogeneous systems.
Jinbang Huang, Yuanzhao Hu, *Zhiyuan Li*, *Ran Qi*, Yixin Xiao, Zhanguang Zhang, Mark Coates, Tongtong Cao, Yingxue Zhang
Huawei Noah’s Ark Lab
cs.RO
Submitted: 2026-07-20
Updated: 2026-09-29
Comments: Best Paper Award at ECCV 2026 Agent in the World Workshop
Code: https://github.com/Physical-Intelligence/openpi
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 88/100
The gist: RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills
Key concepts
- Agentic Skill Modules
- These are independently developed or hand-designed robot policies that RoboHarness wraps. They act as reusable skills for the main planner. Each module has a 'policy card' detailing its specific capabilities and constraints, allowing the system to treat different robots or skills as interchangeable tools.
- Understanding Skills
- These skills interpret raw task inputs to assess how well current states match policy capabilities. They perform five checks: assessing temporal stability of poses, comparing visual observations to training data, matching task instructions to policy descriptions, checking state compatibility with the next policy's region, and evaluating input quality.
- Memory Bridge
- This mechanism manages execution history through a memory bank. It retrieves relevant past experiences based on both text and visual similarity. It then constructs a local support region around retrieved memories and scores candidate states to determine the best target state for the next policy transition.
- Self-Evolution Skills
- These skills allow RoboHarness to adapt online using execution evidence. They include modifying the overall orchestration structure, tuning hyperparameters like retrieval settings, and updating each policy's capability metadata based on real-world successes and failures.
Terminology
Summary
RoboHarness is a unified framework designed for long-horizon robotic tasks that require diverse capabilities, addressing the limitations of existing planning methods which assume homogeneous skills and fixed applicability. It matters because it introduces capability-aware decomposition and routing for heterogeneous robot policies, enabling stable policy handoffs without joint retraining or shared action representations, thereby improving zero-shot long-horizon planning and out-of-distribution robustness.
RoboHarness Framework Overview
RoboHarness is an agentic orchestration framework where a coding agent serves as the high-level planner and router for heterogeneous robot policies. The system encapsulates independently developed or hand-designed policies as reusable agentic skill modules.
These modules are wrapped with auxiliary skills—Understanding, Memory, and Self-Evolution—which provide the necessary information, memory, and adaptation mechanisms for capability-aware planning. This design allows independently trained or hand-designed policies to be combined without joint retraining or a shared action representation.
Understanding Skills
The Understanding skills interpret raw inputs to extract decision-relevant information that reveals the relationship between the current task state and the capability boundaries of the available policies. RoboHarness includes five types of understanding skills:
-
Uncertainty Assessment, which computes
the mean and variance of estimated object poses within a sliding time window to measure temporal stability.
-
Visual-context assessment, which
projects the current observation into a latent space and compares it with visual embeddings of policy-specific training trajectories.
-
Semantic-context assessment, which
compares encoded task instructions and candidate subtasks with the language descriptions associated with each policy.
-
State-policy compatibility assessment, which uses the Memory Bridge to
score the current end-effector pose and joint configuration, indicating whether the robot state lies within the next policy’s in-distribution region.
-
Input-quality assessment, which
evaluates image sharpness, exposure, noise, and task-relevant object visibility to determine whether the current observation is sufficiently reliable.
Memory Skills and Memory Bridge
Memory skills manage execution histories through a memory bank M organized as linked-node trajectories,
where each node stores a subtask instruction, an observation, a robot-state vector, and their corresponding text and visual embeddings. The core mechanism for inter-policy transitions is the Memory Bridge. This bridge combines memory retrieval with spatial distribution learning to preserve spatial consistency.
-
Retrieval: The system performs
hierarchical retrieval over the nodes in M
based on semantic relevance (text similarity) followed by visual relevance (visual embedding similarity). -
Spatial Distribution Construction: Anchor nodes are expanded along their trajectories to construct a local support region, denoted as
Rconf,t = s ∈ R ds d(s, Sret,t) ≤ ϵ,
which restricts the learned progress function to states locally supported by retrieved memory. -
Spatial State Scoring: A lightweight progress estimator predicts the local progress of candidate robot states within Rconf,t; a positive value indicates
greater task progress,
while a negative value indicates less. -
Bridge Trajectory Generation: The system selects the handoff target state as
s∗t = arg max s [fscore,t(s) − λmotionCmotion(st, s)]
subject to constraints ensuring the target lies within the execution distribution and exhibits forward progress relative to retrieved anchors.
Self-Evolution Skills
The Self-evolution skills adapt RoboHarness using online execution evidence through four components:
-
Policy Adaptation: Implements methods like SIMPACT for grasp-pose adjustment and PDDLLM for learning new logical representations online.
-
Harness Refinement: Uses the coding agent to
modify the orchestration structure, including task-routing strategies, skill-invocation rules, and inter-skill coordination.
-
Parameter Tuning: The coding agent predicts movement directions on a search grid to adjust hyperparameters like
retrieval settings, state-scoring thresholds, and bridge-trajectory acceptance criteria.
-
Metadata Update: This component
continually refines each policy’s capability metadata using newly observed successes and failures,
which is used in subsequent planning and policy-routing decisions.
Heterogeneous Policies as Agentic Skills
RoboHarness wraps each heterogeneous robot policy as a callable agentic skill, retaining its native implementation while recording its capabilities in a policy card.
This card records its type, capabilities, interface requirements, constraints, assumptions, training tasks, and historical statistics that summarize policy execution outcomes.
The coding agent can inspect the policy’s implementation code to reason about capabilities and compatibility beyond the summarized metadata. The three underlying policies used are:
-
π0.5: An open-source checkpoint fine-tuned on LIBERO benchmarks for broad language-conditioned manipulation capabilities.
-
OpenVLA-OFT: A checkpoint post-trained with Group Relative Policy Optimization (GRPO) on LIBERO-90 tasks using RLinf reinforcement learning.
Improvements for AI systems
Based on a rigorous review of RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies,
here are specific, high-impact improvements that can be derived from this framework, along with the resulting capabilities of the improved AI system.
The core improvement is transitioning from static or homogeneous policy execution to a dynamic, context-aware orchestration layer capable of handling distribution mismatch
and capability boundaries
between independently developed control systems.
Here are the specific improvements:
-
A Unified, Modular Orchestration Framework (RoboHarness):
-
A Three-Tier Auxiliary Skill Architecture (Understanding, Memory, Evolution Skills):
-
The Implementation of a
Memory Bridge
for Stable Inter-Policy Handoffs: -
Capability-Aware Decomposition and Dynamic Policy Routing:
These improvements will enable the resulting AI system to perform the following specific tasks:
-
Solve complex, long-horizon robotic tasks (e.g., assembling intricate structures or performing multi-step manipulation) that exceed the inherent capabilities of any single, specialized policy (like a Vision-Language model alone or a pure RL controller).
-
Achieve zero-shot generalization to novel task compositions by dynamically selecting and chaining complementary skills from a library of heterogeneous policies (e.g., using a VLA for semantic grounding, TAMP for geometric precision, and RL for reactive contact handling).
-
Exhibit superior robustness under out-of-distribution (OOD) perturbations (e.g., unexpected object placements, noisy sensor data, or imprecise language instructions) by dynamically routing subtasks to the policy whose training distribution best matches the current execution context.
-
Accurately characterize and utilize context-dependent policy capability boundaries, enabling the system to infer when a specific policy is reliable based on real-time observations (e.g., switching from a VLA-dominant mode under good lighting to a TAMP-dominant mode when fine geometric precision is required).
-
Maintain execution consistency across handoffs between policies, even when the terminal state of the preceding policy does not fall within the input distribution of the next policy, by generating
bridge trajectories
guided by retrieved execution memories. -
Continuously improve its own orchestration strategy and underlying policies through an online closed-loop feedback mechanism (Self-Evolution Skills), allowing it to refine task decomposition rules, tune parameter settings for memory retrieval, and adapt individual policies based on accumulated success or failure evidence.
In summary, the improved system is a generalist robotic planner that doesn't just execute tasks; it intelligently manages the coordination of diverse specialized agents to achieve complex goals reliably in unpredictable real-world environments.
Abstract
Long-horizon robotic tasks require a breadth of capabilities beyond what any single existing robot control policy can reliably provide. Combining heterogeneous policies with complementary strengths offers a promising solution, but introduces two key challenges: uncertain capability boundaries and distribution mismatches during policy handoffs. These challenges remain largely unaddressed by existing planning methods, which typically assume homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed heterogeneous policies, including vision-language-action models (VLAs), world-action models (WAMs), reinforcement learning (RL) policies, and task and motion planners (TAMP), as reusable agentic skills. RoboHarness integrates understanding, memory, and evolution skills to reason about policy capabilities and support capability-aware task decomposition and policy routing. To mitigate distribution mismatches during policy handoffs, we introduce Memory Bridge, a plug-in policy-chaining mechanism that enables reliable transitions between heterogeneous policies without joint retraining. Extensive experiments across five public benchmarks, 500 customized tasks across 10 classes, and 135 real-robot trials demonstrate substantial gains in long-horizon and memory-dependent tasks, as well as robustness to out-of-distribution conditions.
Sources
- From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Tree-Planner: Efficient Close-loop Task Planning with Large Language Models
- One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
- H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Adversarial Skill Chaining for Long-Horizon Robot Manipulation via Terminal State Regularization
- Causal World Modeling for Robot Control
- RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
- VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
- Learning to Compose Skills
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving