Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents

arXiv:2607.08448 · cs.RO · Submitted 2026-07-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents".

Rosa: Harness VLA is a memory-augmented agentic framework that exposes a frozen Vision-Language-Action (VLA) model as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives…

Dev: First, who's behind it and why it matters.

Title and authors: Dev: Now that we’ve talked about the high-level concept, let’s look deeper into the actual mechanics of Harness VLA; it really hinges on this division of labor between the LLM planner and the frozen VLA ACT.

Rosa: That division is key because it means we aren't asking that large, expensive VLA to solve every single problem from start to finish; it’s smart enough to know when to step back and let a specific tool handle a difficult moment.

Taro: I see how that works when things get unexpected; if we have an unknown object pose, the planner uses those analytic tools to reposition the robot until it gets into a position where the frozen VLA model can actually find its footing. That’s a clever way to handle uncertainty before even engaging the complex visual-language model.

Dev: And what’s crucial is how they manage that "local" definition; they enforce perception isolation, making sure the VLA only sees what it needs to see for that specific contact interaction, which helps keep the execution loops tight.

Rosa: It really means we’re not asking that large, expensive VLA to solve every single problem from start to finish; we’re just giving it the precise local instructions it needs for the actual touch part.

Taro: That tight loop is interesting because if the world misbehaves mid-action, knowing that only a local segment is relying on the frozen policy means we can isolate that failure to that specific moment in time. It’s like having a safety net for every single grasp.

Dev: That isolation is vital from an engineering standpoint because it helps us keep the execution loops tight, which directly impacts latency during those critical contact phases.

Rosa: It makes the system feel much more controllable, doesn't it? We move from hoping the whole policy works perfectly to knowing that we have reliable building blocks and a good way to stitch them together.

Taro: I think what really stands out is how they use memory—specifically Task Specific Memory—to serialize those successful sequences, which lets the planner re-use that exact strategy later, even if the environment has shifted slightly.

Dev: Yeah, and that ability to re-use successful strategies is what makes it more than just a one-shot execution; it gives the system persistence across different states.

Rosa: That persistence across failures is what’s exciting for autonomy; it means the system doesn't have to relearn everything from scratch when things get messy in an unknown environment.

Taro: I'm eager to see how this approach integrates with other autonomy methods, because having these reliable, structured primitives is a necessary step for any truly long-horizon agent that needs to operate outside of perfect simulation.

Dev: It’s a solid framework for deployment, provided the latency stays in check during those contact-rich phases.

Rosa: It definitely opens up a lot of doors for real-world applications where we need robustness against unexpected shifts in layout or object placement.

The paper's summary: Dev: Moving on to the actual improvements they suggest, it boils down to extending the pretrained VLA model without needing any of that expensive fine-tuning we usually have to do when we want to adapt it for a new task.

Rosa: That's exactly what excites me from an application standpoint; it means we can take a model trained on one set of objects and apply it to another setup just by giving the agent better guidance on how to stage things.

Taro: I think that directly tackles the issue of brittleness; if the VLA struggles with a specific object pose, Harness VLA uses analytic primitives to get it into a configuration where the frozen VLA can take over successfully, which is like building scaffolding for that skill.

Dev: They also suggest that by learning the operating range of these fixed primitives from execution traces, they create a much more robust set of rules for deciding when to use which tool, which is something traditional policy training often struggles with when dealing with noisy data.

Rosa: I’m also really impressed by how they isolate non-contact execution; the analytic primitives handle all the surrounding movement—the transport, the posture adjustments—while VLA ACT just handles the actual contact interaction itself; that separation seems like a very clean way to build reliability.

Taro: That separation makes debugging much clearer too, I think, because you can pinpoint exactly where a failure happened: was it in the planning logic choosing the wrong analytic tool, or was it in the VLA ACT failing at the contact? That level of detail is really useful for research.

Dev: From an engineering standpoint, isolating non-contact execution from contact-rich control means we can better predict timing issues because we know exactly when that handoff to VLA ACT happens.

Rosa: It makes the system feel much more controllable, doesn't it? We move from hoping the whole policy works perfectly to knowing that we have reliable building blocks and a good way to stitch them together.

Taro: I’m eager to see how this approach integrates with other autonomy methods, because having these reliable, structured primitives is a necessary step for any truly long-horizon agent that needs to operate outside of perfect simulation.

Dev: It’s a solid framework for deployment, provided the latency stays in check during those contact-rich phases.

Rosa: It definitely opens up a lot of doors for real-world applications where we need robustness against unexpected shifts in layout or object placement.

The paper's improvements: Rosa: So, looking at "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents," what I'm taking away is that this framework uses memory and fixed primitives to make frozen models much more reliable for complex manipulation tasks.

Dev: It really highlights the importance of that structured decomposition, Rosa, where the planner handles the high-level reasoning while the VLA is just a specialist in contact actions; it’s a smart way to manage complexity.

Taro: I agree with that focus on specialization; when you consider how this handles things when the world throws a curveball, isolating that control loop gives us much better diagnostic capabilities and a clearer path to recovery.

Rosa: It’s fascinating how they use those memory stores—Task Specific Memory for the specific sequence and Global Memory for general success rules—to allow the AI to learn from its own experiences in a way that doesn't require constant retraining.

Dev: From an engineering standpoint, I’m always thinking about the loop rate and latency, and this paper gives us a solid architecture where we can predict those timing issues because we know exactly when the planner hands off control to the VLA ACT.

Taro: That structural knowledge is key for autonomy; it means when things go sideways in a novel environment, the AI has a way to re-frame its approach based on what it already knows about successful sequences.

Rosa: The implication here is that we can deploy powerful pre-trained models into real labs much faster and with more confidence because we’re not trying to teach them everything from scratch for every single new task.

Dev: And the auditability they built in through that JSONL tracing is what really sells it for deployment; we get a clear record of exactly how the AI decided to move and where it made errors.

Taro: That transparency is crucial because we need to understand *why* it failed so we can improve the system's reasoning structure later on.

Rosa: It’s a significant step toward making complex AI agents practical for real-world manipulation tasks, moving them out of just simulation and into actual use cases where things are unpredictable.

Dev: I'm just hoping they can keep those execution loops fast enough for high-frequency contact tasks, because if the loop rate dips too low, all that memory and planning context might become useless.

Taro: That’s a valid concern; the speed of execution is always something we have to watch closely when scaling these agentic systems up.

Rosa: Well, it was a really illuminating look at how we can steer those powerful frozen models into reliable manipulation primitives using memory-guided agents in their paper, "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents."

Dev: It’s a solid framework for deployment, provided the latency stays in check during those contact-rich phases.

Taro: I'm looking forward to seeing how this method helps us tackle more unpredictable, real-world manipulation challenges next.

Conclusion: Rosa: So, to wrap up our discussion on "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents," we’ve established that this framework uses memory and fixed primitives to make large models reliable for manipulation tasks in real-world settings.

Dev: It really highlights the importance of that structured decomposition, Rosa, where the planner handles high-level reasoning while the VLA is just a specialist in contact actions; it’s a smart way to manage complexity.

Taro: I agree with that focus on specialization; when you consider how it handles things when the world throws a curveball, isolating that control loop gives us much better diagnostic capabilities and a clearer path to recovery.

Rosa: It’s fascinating how they use those memory stores—Task Specific Memory for the specific sequence and Global Memory for general success rules—to allow the AI to learn from its own experiences in a way that doesn't require constant retraining.

Dev: From an engineering standpoint, I’m always thinking about the loop rate and latency, and this paper gives us a solid architecture where we can predict those timing issues because we know exactly when the planner hands off control to the VLA ACT.

Taro: That structural knowledge is key for autonomy; it means when things go sideways in a novel environment, the AI has a way to re-frame its approach based on what it already knows about successful sequences.

Rosa: The implication here is that we can deploy powerful pre-trained models into real labs much faster and with more confidence because we’re not trying to teach them everything from scratch for every single new task.

Dev: And the auditability they built in through that JSONL tracing is what really sells it for deployment; we get a clear record of exactly how the AI decided to move and where it made errors.

Taro: That transparency is crucial because we need to understand *why* it failed so we can improve the system's reasoning structure later on.

Rosa: It’s a significant step toward making complex AI agents practical for real-world manipulation tasks, moving them out of just simulation and into actual use cases where things are unpredictable.

Dev: I'm just hoping they can keep those execution loops fast enough for high-frequency contact tasks, because if the loop rate dips too low, all that memory and planning context might become useless.

Taro: That’s a valid concern; the speed of execution is always something we have to watch closely when scaling these agentic systems up.

Rosa: Well, it was a really illuminating look at how we can steer those powerful frozen models into reliable manipulation primitives using memory-guided agents in their paper, "Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents."

Dev: It’s a solid framework for deployment, provided the latency stays in check during those contact-rich phases.

Taro: I'm looking forward to seeing how this method helps us tackle more unpredictable, real-world manipulation challenges next.

Tsinghua University · Purdue University · Institute of Automation, Chinese Academy of Sciences · Infinigence AI · Hong Kong University of Science and Technology · Zhongguancun Academy

cs.RO

Submitted: 2026-07-09

Updated: 2026-09-24

Code: https://github.com/RLinf/RPent

Project page: https://harnessvla.github.io

Importance score: 92/100

The gist: Harness VLA is a memory-augmented agentic framework that exposes a frozen Vision-Language-Action (VLA) model as a retryable contact-rich primitive and composes it with a small fixed library of

Key concepts

Harness VLA
A memory-augmented agentic framework that exposes a frozen Vision-Language-Action (VLA) model as a retryable contact-rich primitive. It composes this frozen model with a small fixed library of analytic primitives to create reliable manipulation capabilities.
Division of Labor
The key mechanic where the LLM planner handles high-level reasoning, while the frozen VLA acts as a specialist for contact actions. This allows the system to step back and let a specific tool manage difficult moments instead of asking the large VLA to solve everything.
Task Specific Memory
A memory store used by Harness VLA to serialize successful sequences. This allows the planner to re-use exact strategies later, even if the environment has slightly shifted, providing persistence across different states and failures.
Perception Isolation
Enforcing that the VLA only sees what it needs to see for a specific contact interaction. This keeps execution loops tight and helps ensure the VLA model only focuses on the necessary local visual information.

Terminology

Summary

Harness VLA is a memory-augmented agentic framework that exposes a frozen Vision-Language-Action (VLA) model as a retryable contact-rich primitive and composes it with a small fixed library of analytic primitives for grounding, staging, transport, navigation, and release. Rather than expanding the skill library, the harness learns the operating range of these fixed primitives from task-specific execution traces, global success rules, and failure models. By lifting semantic re-grounding, non-contact execution, and VLA re-staging to the planner while reserving the frozen VLA for local contact-rich phases, Harness VLA extends pretrained VLAs beyond their original trajectory distribution without fine-tuning.

The framework follows an agentic execution loop where a high-level agentic planner reasons over a fixed primitive library and retrieves context from Task Specific Memory and Global Memory. The primitive library includes analytic primitives (like MOVE TO, ROTATE, SET GRIPPER) for perception-conditioned staging, transport, posture adjustment, and release; a structurally special VLA primitive called VLA ACT for contact-rich behaviors; and mobile-base primitives (like NAVIGATE TO, MOVE BASE) for kitchen-scale staging.

The framework operates in two phases:

  1. Exploratory Bootstrapping Phase: The agent autonomously interacts with the environment to discover a working solution, focusing on iterative composition of the learned VLA primitive and deterministic analytic primitives. Upon success, it abstracts its experience into Task Specific Memory (a serialized JSONL trace and a semantic JSON summary) and Global Memory (reusable success rules and failure models).

  2. Deployment Evaluation Phase: During formal evaluation on unseen environment variations, the agent retrieves the pre-computed JSONL trace from Task Specific Memory and grounds it dynamically using live RGB-D observation, while referencing success rules and failure models from Global Memory to execute the trajectory deterministically.

The core contributions are:

"A memory-augmented agentic framework for using a frozen VLA as a primitive. Harness VLA composes VLA ACT with fixed analytic primitives, extending a pretrained VLA from local contact-rich control to long-horizon, perturbed manipulation without fine-tuning the VLA or expanding the primitive vocabulary at deployment time."

"An empirical analysis showing why a small fixed primitive library is sufficient when the planner learns how to use it. Repeated planner-staged invocations can reframe brittle VLA attempts, while analytic primitives solve much of the non-contact structure around each contact-rich phase."

The framework utilizes a file-mediated Read-Eval-Print Loop (REPL) protocol where the planner emits structured JSON commands to a worker that executes the selected primitive in the live environment, refreshes observations, and logs diagnostic records. The prompt given to the agent is structured into several modules: Role and Success Signal, Perception Isolation (enforcing grounding from RGB/depth maps), File-Based REPL control, Primitive Vocabulary specification, Division of Labor between Agent and VLA (using VLA ACT for contact-rich phases and analytic primitives for grounding/transport), Task Language authority, Seed 0 Task Specific Memory retrieval (using the JSON audit to understand strategy and the JSONL trace to recover primitive order), Global Memory utilization (for success rules and failure models), Closed-Loop Verification and Recovery, Budget/Reset policy, Output Discipline (writing audit JSON and command-trace JSONL), and an Operating Loop.

The analysis reveals three key findings:

"Key Finding 1: Planner-level semantic re-grounding restores task-conditioned behavior. The massive gap between Harness VLA and end-to-end VLAs in Table 3 is achieved without altering the visuomotor backbone... Harness VLA makes semantic grounding explicit at the planner level: the planner Π parses the task description, resolves the current contact target from the live RGB-D observation, uses analytic primitives for staging and repositioning, and invokes or re-invokes VLA ACT only for the local contact-rich phase (Key Finding 2)."

"Key Finding 2: Planner-staged VLA invocation improves frozen-policy reliability. In Harness VLA, the planner does not call the VLA as a one-shot black box... The planner Π therefore treats VLA ACT as a local contact-rich primitive whose invocation can be re-staged and retried."

"Key Finding 3: Analytic primitives isolate non-contact execution from contact-rich control. Analytic primitives do not replace the VLA on contact-rich operations. Instead, they handle the surrounding noncontact structure of the task: free-space transport, pre-contact staging, wrist or base reorientation, retreat, and post-contact repositioning."

Empirical results show significant improvements over strong baselines: Harness VLA improves over LIBERO-Pro by 38.6 and RoboCasa365 by 27.1 percentage points on LIBERO-Pro and RoboCasa365, respectively, reaching 58.4% on RoboTwin C2R. The usage statistics confirm the asymmetric decomposition: in LIBERO, analytic primitives dominate (MOVE TO accounts for 61.8%), while in RoboCasa365 and RoboTwin C2R, VLA ACT rises to a dominant share (35.3% and 47.4%, respectively), reflecting the division of labor between analytic primitives handling reproducible geometry/staging and VLA ACT supplying learned local interactions.

The framework is evaluated across four benchmark families: LIBERO, LIBERO-Pro, RoboCasa365, and RoboTwin C2R. The performance is assessed in both few-shot (with Task Specific Memory) and zero-shot regimes (without memory retrieval). For example, on RoboTwin C2R under clean-to-randomized transfer, Harness VLA achieves a 58.4% success rate with the LingBot-VLA backend. The system is designed to be reusable across different manipulation benchmarks by using a shared prompt core and abstracting heterogeneous VLA models as the single contact-rich primitive VLA ACT. The framework is limited by an open feedback loop between the high-level planner and low-level VLA, and lacks joint fine-tuning via environmental rewards. A future direction involves combining this fixed-vocabulary composition strategy with automatic skill-discovery systems like ASPIRE to propose new reusable skills while retaining the auditable primitive interface.

The evaluation protocol ensures auditability by storing a procedural JSONL trace in Task Specific Memory and distilling success rules into Global Memory, allowing the planner to transfer structural knowledge without replaying literal coordinates. The agent is explicitly prohibited from accessing privileged simulator state or ground-truth object poses, enforcing realistic partial-observation control. The final output discipline requires writing both an audit JSON and a command-trace JSONL for every rollout.

The primitive vocabulary is fixed before evaluation, meaning the planner cannot invent new primitives at deployment time, which supports the claim that a small fixed primitive library is sufficient when the planner learns how to use it. The system demonstrates robustness against semantic re-targeting and spatial-layout shifts by leveraging analytic primitives to handle non-contact structure around local contact-rich phases. The overall framework successfully extends pretrained VLAs beyond their original trajectory distribution without fine-tuning or deployment-time primitive expansion.

The final summary of the division of labor is: analytic primitives organize the task around contact-rich phases, while the VLA remains responsible for the phases where learned visuomotor control is needed." (Page 39)

The system uses a shared prompt core that defines its responsibilities and closed-loop behavior across all benchmarks, while benchmark-specific prompts fill in environment-context fields like success signals, frozen VLA entry points, and required audit artifacts. This ensures the framework is reusable for additional manipulation benchmarks. The structure of the agent's lifecycle is defined by this shared prompt core and the specific environmental assumptions provided by the benchmark-specific instantiations." (Page 36)

The system enforces perception isolation through a rule that requires identifying objects from RGB/depth maps and indexing them into a world map, explicitly prohibiting access to hidden state variables or simulator internals." (Page 37)

The framework's success is validated by strong benchmark results across standard and perturbed tabletop manipulation, household kitchen manipulation, and clean-to-randomized transfer." (Page 4)

The system leverages memory for both procedural context (Task Specific Memory: JSONL trace) and task-independent knowledge (Global Memory: success rules/failure models), refined iteratively rather than simply accumulated." (Page 34)

The system's core mechanism is the delegation of semantic grounding, spatial decomposition, and long-horizon task planning to the LLM planner, while each VLA is invoked solely for localized interaction execution conditioned on the current observation." (Page 31)

The system's primary limitation is an open feedback loop between the high-level planner and low-level VLA, and its lack of joint fine-tuning via environmental rewards and human preferences." (Page 16)

The framework's success is achieved by surrounding a frozen policy with an auditable execution loop, fixed primitive contracts, memory, feedback, and task-level verification." (Page 15)

The system's final operational loop is: Read prompt, state, task language, perception, Task Specific Memory, and Global Memory. Localize entities. Execute one primitive. Observe. Recover. Repeat. (Page 38)

This structure ensures that pretrained VLAs are most effective when isolated to contact-rich visuomotor control; abstracting semantic and spatial bindings away from the VLA prevents the catastrophic failures frequently observed in monolithic deployments." (Page 15)

The system is designed to be reusable across different manipulation benchmarks by using a shared prompt core and abstracting heterogeneous VLA models as the single contact-rich primitive VLA ACT." (Page 32)

The system's final operational loop is: Write one JSON command to command.json. Wait until the driver finishes that primitive, read the new state NN.json, log NN.json, images, depth maps, and world maps. Then decide the next command. (Page 37)

The system's final operational loop is: Read prompt, state, task language, perception files, Task Specific Memory and Global Memory. Localize entities. Execute one primitive. Observe. Recover. Repeat. (Page 38)

The system's final operational loop is: Write one JSON command to command.json; wait for the driver to finish that primitive; read the new state NN.json, log NN.json, images, depth maps, and world maps. (Page 37)

The system's final operational loop is: Read prompt, state, task language, perception files and Task Specific Memory and Global Memory. Localize entities. Execute one primitive. Observe. Recover. (Page 38)

The system's final operational loop is: Read prompt, state, task language, perception files and Task Specific Memory and Global Memory. Localize entities. (Page 38)

The system's final operational loop is: Read prompt, state NN.json; log NN.json; images; depth maps; world maps. (Page 37)

The system's final operational loop is: Read prompt, state NN.json; log NN.json; images; depth maps. (Page 37)

The system's final operational loop is: Read prompt, state NN.json; log NN.json; images. (Page 37)

The system's final operational loop is: Read prompt, state NN.json; log NN.json. (Page 37)

The system's final operational loop is: Read prompt, state NN.json; log NN.json. (Page 37

Improvements for AI systems

Here are specific improvements to AI systems based on the Harness VLA framework, and what these improved systems can achieve:


) 1. Robustness to Distribution Shifts (Perturbation Handling):

The system can reliably execute manipulation tasks even when the environment changes significantly from its training data (e.g., semantic retargeting or spatial-layout shifts).

  • Improved System Capability: The agent decomposes the task into analytic primitives for non-contact structure (navigation, staging, re-orientation) and invokes the frozen VLA only for local contact phases. This prevents the monolithic VLA from failing catastrophically when semantic targets or layouts shift.

) 2. Long-Horizon Composition with Reusable Skills:

The system can solve complex, multi-step tasks by composing fixed analytic controllers and the learned VLA primitive intelligently, rather than relying on an end-to-end policy that must learn every sequence from scratch.

  • Improved System Capability: By learning operating ranges of fixed primitives via Task Specific Memory and Global Memory, the planner can dynamically re-stage the robot around a new target or re-attempt a failed contact phase using analytic control, effectively extending the VLA's utility beyond its original trajectory distribution without requiring further fine-tuning.

) 3. Efficient Deployment with Frozen Policies:

The system allows deployment of highly capable, pre-trained Vision-Language-Action (VLA) models in real-world settings without the massive computational cost and risk of fine-tuning them on new tasks.

  • Improved System Capability: The VLA is treated as a contact specialist (VLA ACT), specialized for local, contact-rich control. The planner handles all high-level reasoning (language grounding, navigation), drastically reducing the need to adapt or retrain the core visuomotor policy for every new task variant.

) 4. Auditable and Transparent Decision Making:

The system provides a complete record of its reasoning process, allowing researchers to diagnose failures and understand success strategies post-hoc.

  • Improved System Capability: The file-mediated REPL protocol mandates writing an audit JSON and a command JSONL after every primitive invocation. This trace explicitly records the order of analytic vs. VLA calls, transition points between phases (staging, contact, transport), and failure diagnoses, making the agent's behavior fully inspectable and reproducible.

) 5. Adaptive VLA Invocation for Contact-Rich Tasks:

The system optimizes the use of expensive VLA computations by only invoking them when necessary for local interaction stability.

  • Improved System Capability: The planner uses analytic primitives to bring the robot into a favorable VLA-compatible local observation (re-staging). Only then is VLA ACT invoked for a short, localized burst of contact manipulation, and the system can re-stage or retry if the contact outcome is unstable, ensuring sparse but highly effective VLA utilization.

) 6. Scalability Across Diverse Embodiments:

The framework supports heterogeneous robotic setups (e.g., single-arm tabletop vs. mobile base kitchen robots vs. dual-arm bimanual systems).

  • Improved System Capability: The fixed primitive vocabulary (Table 8) and the unified VLA ACT interface allow the same high-level agentic planner to operate across different robot hardware, provided the environment exposes the necessary primitives (e.g., mobile base primitives for kitchen tasks or dual-arm binding for bimanual tasks).

) 7. Task-Specific Knowledge Transfer:

The system can rapidly adapt to new tasks by leveraging past successes without requiring a full retraining cycle.

  • Improved System Capability: The Task Specific Memory stores the successful primitive sequence (the procedural skeleton) from a reference seed. When deployed on a new task, the planner reuses this structure, grounding only the spatial arguments in the current observation, enabling few-shot re-grounding for immediate performance gains.

Sources

Related papers