Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness

arXiv:2606.18363 · cs.RO, cs.AI · Submitted 2026-06-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness".

Dev: Guava presents a harness framework for embodied manipulation that identifies three key ingredients for effective embodied agents: iterative reasoning, semantic action abstractions, and multimodal observations.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we've been looking at the paper, "Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness." It seems like the core idea is building this harness to make large vision-language models capable of actual physical manipulation without needing massive amounts of training data.

Dev: That's right, Rosa; it’s about creating a structured interface that lets these big models use external tools for perception and control instead of trying to learn everything internally. It’s a shift from end-to-end systems to something more modular, which I think is important for deployment reliability.

Taro: I'm curious how this structure handles the inevitable failures when the agent interacts with the real world, Rosa. Does this harness give it enough tools to recover when things go sideways?

Rosa: That’s a big question, Taro; one of the main points they make is that effective agents need iterative perception-reasoning-action loops to handle those execution outcomes and recover from failures. This closed loop is supposed to allow the AI to adapt its plan as it goes.

Dev: I agree with Rosa; having that feedback mechanism built in is crucial for any practical system, especially when you're dealing with physical tasks where things don't always go exactly as planned. The paper suggests these loops are essential for recovering from grasp failures or state deviations during manipulation.

Taro: If the system encounters something completely unexpected, like an object moving unexpectedly or a grasp slipping, how does this iterative loop manage that deviation? Does it have a specific mechanism for re-planning?

Rosa: The paper points to semantic action abstractions as another key ingredient; these allow the language model to focus on high-level planning rather than getting bogged down in all the low-level geometric and physical reasoning. They use tools like "grasp(object)" or "align(object, direction, clearance)" instead of forcing the VLM to handle every tiny detail itself.

Dev: Focusing on high-level planning via those semantic abstractions sounds much more manageable for a model with fewer training examples than trying to teach it raw motor control. It seems like they're reducing the burden on the language model significantly.

Taro: Reducing that burden is smart, but what about how it handles the visual and textual information simultaneously? The paper emphasizes multimodal observations, combining visual input with text representations for better grounding during sequential decisions. Is that enough to make things truly robust?

Title and authors: Rosa: They argue that combining visual observations with textual state representations improves grounding and reduces ambiguity when the agent is making a series of decisions in sequence. This combination helps the AI understand what it's seeing in relation to what it's trying to do.

Dev: I see how that helps with sequential decision-making; having both streams of information gives the model richer context for each step, which should lead to more stable control signals and less jitter in the execution. But we still need to think about the latency introduced by all those perception steps.

Taro: Thinking about latency is key, Dev; if the loop rate isn't fast enough, even a perfect plan will fail in a dynamic environment where things are changing rapidly. Does this harness design inherently allow for high-frequency updates necessary for real-time interaction?

Rosa: The framework itself is designed to encourage embodied reasoning and tool-calling for manipulation, which implies that the structure supports efficient interaction between the language model's high-level plans and the low-level tools. It’s about structuring that communication effectively.

Dev: Structuring it well is only half the battle; we have to ensure those tool calls happen fast enough so we aren't waiting around for perception outputs that are too slow for dynamic tasks. The paper focuses more on the logic flow than the hardware timing constraints, which is something I’d want to see addressed more directly in future work.

Taro: And what about generalizing this setup? When we move from simulation to a real-world setting, how reliable is this distilled capability? Can we trust that the learned harness works as well outside of the training environment?

Rosa: They demonstrate that this approach enables compact open-source models to acquire strong manipulation capabilities with minimal training data, and they've shown success in transferring these skills from simulation to the real world without needing additional real-world fine-tuning.

Dev: That transfer capability is impressive, Rosa; if the distillation pipeline works as described, it suggests that the core logic learned through those structured interactions generalizes well across different model architectures. I’m interested in how they handled that data efficiency aspect when distilling into a 4B model using fewer than 2K trajectories.

Taro: That low data requirement is what makes this scalable for smaller deployments, but it raises questions about the depth of knowledge the distilled agent actually possesses compared to a massive proprietary model. What's the trade-off there?

Title and authors: Rosa: The results show that Guava-Agent-4B achieves performance comparable to frontier proprietary models across diverse evaluation scenarios, including out-of-distribution tasks, which suggests that for manipulation specifically, this distilled approach yields strong performance.

Dev: Comparable to GPT-five point four at seventy point two percent and CaP-Agent0 at sixty-two point seven percent is a significant number when you consider the model size and the fact that it was trained on such limited simulation data—it shows the harness structure itself is powerful enough to extract useful skills from sparse inputs.

Taro: I'm still focused on what happens when things misbehave; if the model encounters an object it hasn't seen before, does this harness rely too heavily on its pre-programmed semantic tools, or can it devise a completely novel way to manipulate that new thing?

Rosa: The paper shows strong generalization to unseen objects and prompts, achieving one hundred percent success on OOD object tasks like picking up a carrot or a lemon in a bin. This indicates that the semantic abstraction allows for good compositionality, letting the agent adapt its plan even when encountering novel physical entities.

Dev: That OOD success is compelling because it means the system isn't just memorizing specific trajectories; it’s using the harness to reason about the underlying affordances of whatever new object it sees, which is a much more robust form of capability. I wonder if that reasoning complexity impacts our latency metrics.

Taro: It does impact complexity, Dev; but the structure seems designed to manage that complexity through decomposition rather than letting it crash the execution loop. The authors also showed that reinforcement learning post-training can substantially improve long-horizon reasoning and recovery behaviors, pushing performance on tasks like "shell game" from six point seven percent up to sixty point zero percent.

Rosa: That RL post-training result is really telling, Taro; it shows that while the harness provides a strong foundation, optimizing the policy further through reinforcement learning really sharpens those long-horizon recovery skills and makes the agent much more capable in complex scenarios.

Dev: From an engineering standpoint, that sixty point zero percent improvement on a challenging task like "shell game" is exactly what we need to see if we want this to be a reliable system for more complex, real-world applications where failure recovery is non-negotiable. It shows the RL component really helps solidify those recovery behaviors the harness sets up.

Title and authors: Taro: So, to wrap up on the methodology, they distill these capabilities into a 4B open-source model using under 2K trajectories collected in simulation, and this distillation pipeline involves generating both successful and recovery trajectories from perturbed states. This data-efficient distillation pipeline is central to making it accessible.

Rosa: Exactly; that whole process—from data generation engine with scene randomization to processing those recovery trajectories—is what makes the Guava framework a viable interface for widely available, compact models. It turns frontier VLM capabilities into something smaller and more deployable.

Dev: I just want to stress the practical implication: we get high performance, strong generalization, and real-world transferability without needing massive proprietary datasets for fine-tuning every time we want to adapt a model for a new physical task. That’s a huge operational win.

Taro: And from my side, the implication is that by separating the low-level control from the high-level reasoning via these semantic tools, we are building agents that are more flexible and less brittle when faced with unpredictable real-world situations. It makes them more resilient to misbehavior.

Rosa: So, to summarize, this paper on Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness shows how combining iterative reasoning loops, semantic action abstractions, and multimodal observations creates an effective harness for embodied manipulation.

Dev: And the practical result is that you can distill those skills into a compact model that performs competitively against much larger systems on both in-distribution and out-of-distribution tasks.

Taro: It points toward a future where deploying complex agentic manipulation systems doesn't require training models on millions of physical interactions, but rather structuring the interaction around proven reasoning principles.

Rosa: That’s a lot to take in, Taro; it really suggests that for embodied agents, the way we structure their interaction with tools might be more important than just the raw size of the underlying model.

Dev: I think that’s the core message; it's about designing an interface that is robust and modular enough to handle the messy reality of physical tasks effectively without needing massive amounts of data to learn every single nuance.

Taro: Indeed, and I think this work on Guava opens up a path for developing more adaptable and resilient embodied agents across many different manipulation domains.

Rosa: It certainly gives us a solid foundation for thinking about how to design these interaction strategies moving forward, and it’s an interesting direction to watch in the field.

The paper's summary: Rosa: So, we're looking at the summary of Guava, which boils down to how they've built this harness to take those massive vision-language models and shrink them down into a compact agent that can actually do physical tasks with minimal training data.

Dev: That’s right; essentially, it’s about creating a structured interface that lets these large models use external tools for perception and control instead of trying to learn everything internally, which is a really smart way to manage complexity.

Taro: I'm thinking about the implications for real-world deployment; if you can get strong manipulation capabilities from a much smaller model, does that mean we can actually put these agents in more complex, messy environments without needing massive datasets?

Rosa: They show that by focusing on iterative reasoning and semantic tools, this distilled agent performs quite well across different scenarios, even handling things it hasn't seen before.

Dev: The key result is the efficiency; they managed to distill those manipulation skills into a 4B model using fewer than 2K trajectories collected entirely in simulation, which really shows how powerful that harness structure can be.

Taro: And that data efficiency is something I’m interested in because it lowers the barrier for researchers who don't have access to huge physical interaction datasets to train models.

Rosa: Exactly; this approach means we can create agents that are compact and capable right out of the box, which is a big step toward making embodied AI more accessible.

Dev: From an engineering standpoint, the fact that they achieved performance levels comparable to frontier proprietary models across in-distribution and out-of-distribution tasks is genuinely impressive when you consider the training budget.

Taro: I'm still focused on the real world aspect; how robust is this distilled agent when it steps outside of simulation and actually has to deal with unpredictable physical dynamics?

Rosa: They demonstrated strong transfer from simulation to the real world without needing any additional real-world fine-tuning, which suggests the semantic planning part of the harness is quite effective at separating high-level logic from low-level physical details.

Dev: That sim-to-real transfer capability is a huge deal for deployment; if we can rely on that, it cuts down significantly on the expensive and time-consuming process of fine-tuning every model for every new task.

Taro: So, the implication is that by structuring the interaction this way, we aren't just teaching a model specific movements; we're giving it a framework for reasoning about manipulation itself.

Rosa: That’s right; it shifts the focus from brute-force learning to designing an effective communication structure between a language model and its physical tools.

Dev: It really highlights that for embodied agents, the way we organize their interaction with external capabilities might be more important than just the raw size of the underlying model.

Taro: That leads me to wonder what happens when things go wrong in a dynamic setting; how does this structure allow for recovery beyond just executing a planned sequence?

Rosa: They showed that adding reinforcement learning post-training significantly improved long-horizon reasoning and recovery behaviors, moving performance on difficult tasks up substantially.

Dev: That RL component is crucial; it shows that while the harness provides the right tools for planning, optimizing the policy further through reinforcement learning really sharpens those failure recovery skills.

Taro: It suggests that combining a good harness structure with post-training optimization gives us agents that are not just planners but actually resilient when things inevitably misbehave.

Rosa: So, we're looking at a system where you use the harness to get the agent running efficiently, and then you use RL to make sure it can actually handle the unexpected bumps in the road.

Dev: That combination gives us a very solid foundation for developing agents that are both smart enough to plan and tough enough to recover in physical tasks.

Taro: It really points toward a future where we can deploy complex agentic systems more reliably because we’ve solved part of the data acquisition bottleneck.

The paper's improvements: Rosa: So, to wrap up on the methodology, Guava introduces these three core design principles—iterative loops, semantic abstractions, and multimodal observations—and then they build a framework around them for embodied tool use.

Dev: Right; it’s not just a collection of tools; it's about integrating those concepts into a unified architecture that encourages both embodied reasoning and tool-calling for manipulation.

Taro: I'm thinking about the iterative loop again; if the system encounters an unexpected physical state, how does this framework ensure the agent doesn't get stuck in a loop of failed attempts?

Rosa: The iterative ReAct style loops are specifically designed to let the AI adapt to execution outcomes and recover from failures by having it operate in a closed-loop process.

Dev: That feedback mechanism is key for stability; it means when things go sideways, the agent has a clear path to re-plan rather than just crashing or repeating the same mistake.

Taro: And what about those semantic abstractions they mentioned earlier; how do those tools help the AI manage the complexity of physical reasoning?

Rosa: These abstractions allow language models to focus on high-level planning, significantly reducing the low-level geometric and physical reasoning burden placed on them by providing tools like 'grasp(object)' or 'align(object, direction, clearance)'.

Dev: That’s a major win for model efficiency; it lets the language model handle the strategy while specialized functions manage the precise motor commands.

Taro: I also want to talk about those multimodal observations; how does combining visual and textual state representations actually help in decision-making during sequential actions?

Rosa: The combination of visual observations with textual state representations helps improve grounding and reduces ambiguity when the agent is making a series of decisions in sequence because it gives the AI richer context.

Dev: That context richness should lead to much more stable control signals, which is what we need for reliable execution, even if the loop rate itself has to be managed carefully.

Taro: If we look at the overall improvements they suggest, it seems like they are really pushing toward creating systems that can generalize well and recover effectively in novel situations.

Rosa: They’ve shown that this approach enables compact open-source models to acquire strong manipulation capabilities with minimal training data, which is a big deal for accessibility.

Dev: And the pipeline they developed for distillation shows they can achieve this without needing millions of trajectories; it uses under 2K in simulation and then augments them with recovery trajectories from perturbed states.

Taro: That data-efficient distillation pipeline is very promising because it means we don't need to spend all that time collecting real-world interaction data just to get a basic manipulation skill set.

Rosa: So, the implication here is that we can build these powerful manipulation agents using much smaller models than previously thought possible, and they transfer those skills reliably to the real world.

Dev: It really shows that the harness structure itself is robust enough to extract useful skills from sparse inputs, which makes deployment much more feasible in terms of computational cost.

Taro: I'm still curious about where this work stops; do you see any limitations in this approach when dealing with extremely fast or highly dynamic physical interactions that might break the loop rate assumptions?

Rosa: The authors flag that while the framework is scalable, it’s a harness architecture, so its performance in real-time systems depends heavily on how well we tune those semantic tools to match the dynamics of the environment.

Conclusion: Rosa: So, we're wrapping up our discussion on "Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness," summarizing how this harness framework lets us distill those massive models into compact agents with minimal training data.

Dev: That’s right; the core takeaway is that we can create agents that are performant and deployable by structuring the interaction around proven reasoning principles rather than just relying on raw model size.

Taro: I think the real world impact here is making sophisticated manipulation skills available to a much wider range of researchers who don't have access to massive physical interaction datasets.

Rosa: Exactly; this moves embodied AI toward being more accessible because we can achieve strong capabilities without needing millions of physical trajectories for every new model we want to deploy.

Dev: From an engineering view, the ability to transfer those skills from simulation directly into real-world tasks without extensive fine-tuning is a massive operational win for robotics development.

Taro: I just think the focus on iterative loops and semantic tools means these agents are much more resilient when they encounter unexpected physical behavior in unpredictable environments.

Rosa: That’s true; by building that structure in, we're not just teaching a model to follow a script; we're giving it the ability to adapt its strategy as things change.

Dev: We should keep an eye on how they handle those latency issues, though; if the loop rate isn't fast enough for truly dynamic physical interactions, even the best plan won't execute reliably.

Taro: That’s a valid point, Dev; while the structure is sound, we need to make sure that when things misbehave in real-time, the recovery mechanism actually keeps up with the pace of change.

Rosa: Well said; it’s all about finding that sweet spot between high-level planning and low-level execution speed.

Dev: I think looking at this paper on "Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness" really sets a new benchmark for how we structure agent interaction with tools.

Taro: It certainly does, and it makes me wonder what the next frontier is; are we going to see these harnessed models deployed in truly unstructured environments, or will they stay focused on controlled lab settings?

Rosa: That’s the million-dollar question for field robotics; I’m excited to see if we can take this harness out of the lab and into more complex, uncontrolled settings.

Dev: We need to check those real-world loop rates closely, Rosa; if we want these agents doing anything meaningful outside a controlled environment, that latency has to be tight.

Haowen Liu, *Co-first Author*, Xirui Li, *Co-first Author*, Shaoxiong Yao, Peng Shi, Tianyi Zhou, Jia-Bin Huang, *Co-first Author*, Furong Huang

University of Maryland College Park · University of Illinois Urbana-Champaign · University of Waterloo · Mohamed bin Zayed University of Artificial Intelligence · University of Pennsylvania

cs.RO, cs.AI

Submitted: 2026-06-16

Updated: 2026-09-29

Code: https://github.com/karpathy/autoresearch

Project page: https://guava-harness.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Guava presents a harness framework for embodied manipulation that identifies three key ingredients for effective embodied agents: iterative reasoning, semantic action abstractions, and multimodal

Key concepts

Iterative perception-reasoning-action loops
This involves a closed loop where the agent perceives its environment, reasons about the next step, and then takes an action. This cycle is repeated until a goal is achieved or a failure occurs. It helps the agent adapt to unexpected outcomes, like failing to grasp an object, allowing it to recover and try again.
Semantic action abstractions
These are high-level commands that describe manipulation goals rather than low-level motor movements. Instead of learning complex physical controls, the model learns to use tools like 'grasp(object)' or 'align(object, direction)'. This shifts the focus from precise physical control to strategic planning.
Multimodal observations
This concept combines visual input (what the agent sees) with textual representations (what the agent knows about its state). By using both simultaneously, the system gains a richer understanding of its surroundings. This helps reduce confusion when making sequential decisions and improves how accurately it grounds its actions in reality.

Terminology

Summary

Guava presents a harness framework for embodied manipulation that identifies three key ingredients for effective embodied agents: iterative reasoning, semantic action abstractions, and multimodal observations. This work demonstrates that a well-designed harness can serve as a scalable, model-agnostic interface for embodied manipulation, enabling compact open-source models to acquire strong manipulation capabilities with minimal training data.

Key Ingredients for Effective Embodied Agents

The study identifies three core design principles critical for effective embodied agents:

  1. Iterative perception-reasoning-action loops: iterative ReAct (Yao et al., 2023) loops are essential for adapting to execution outcomes and recovering from failures. This closed-loop reasoning process allows the VLM to operate in a loop that supports recovery from grasp failures and state deviations.

  2. Semantic action abstractions: These allow language models to focus on high-level planning rather than low-level control, as they reduce the low-level geometric and physical reasoning burden placed on the VLM. The framework utilizes semantic manipulation tools like grasp(object) and align(object, direction, clearance).

  3. Multimodal observations: The system combines visual observations with textual state representations to provide complementary information. This combination is noted as improving grounding and reducing ambiguity during sequential decision-making.

Guava Framework Design

Guava is a harness framework for embodied tool use that integrates these principles into a unified agent architecture. It defines structured interaction strategies between an embodied agent and its environment, encouraging embodied reasoning and tool-calling for manipulation. The framework provides a set of semantic tools, such as:

((

grasp(object) Pick up an object using perception-guided grasping.

align(object,direction,clearance) Align the gripper to specified direction around a target object at clearance distance.

get position(object) Query the 3D position of an object.

get position size(object) Query both object position and bounding box size.

move(x,y,z) Move the end effector to a target position.

rotate(angle, axis) Rotate the gripper by angle around axis.

close gripper

release

home pose

Data-Efficient Distillation Pipeline

To investigate whether an effective harness can serve as an universal interface for embodied manipulation across models, the authors developed an end-to-end training pipeline that distills embodied manipulation capabilities into a 4B open-source model using fewer than 2K trajectories collected entirely in simulation. This pipeline involves:

  1. Data Generation Engine: This engine collects interaction trajectories from frontier VLMs operating under the Guava harness, incorporating scene randomization and targeted perturbations to produce both successful and recovery trajectories.

  2. Data Processing: Trajectories are filtered by categorizing execution outcomes, retaining only successful episodes initially, and then augmenting them with recovery trajectories generated from perturbed execution states.

  3. Training Pipeline: The model undergoes a two-stage process: first, supervised finetuning (SFT) on the embodied data engine trajectories to learn manipulation skills and corrective behaviors; second, Group Relative Policy Optimization (GRPO) with a sparse task-success reward.

Performance and Generalization Results

Experimental results show that the resulting model, Guava-Agent-4B, achieves performance comparable to frontier proprietary models across diverse evaluation scenarios. Key findings include:

: Guava-Agent-4B achieves the strongest overall performance across ID and OOD tasks. It outperforms GPT-5.4 (70.2%) and CaP-Agent0 (62.7%). On In-Distribution (ID) tasks, it consistently achieves the best performance, including 100% success on place can in box and remove cube from tray. Furthermore, it exhibits strong generalization to unseen objects and prompts. 100% success is achieved on OOD object tasks like pick up carrot and lemon in bin. 93.3% success is achieved on long-horizon tasks such as separate food and utensils and set table.

**: The model demonstrates effective transfer from simulation to the real world, achieving the highest overall success rates on both ID (86%) and OOD (92%) real-world tasks. This transfer occurs without additional real-world fine-tuning. **

**: Reinforcement Learning post-training substantially improves long-horizon reasoning and recovery behaviors. Comparing SFT and RL versions, RL post-training improves performance on challenging tasks like shell game from 6.7% (SFT) to 60.0% (RL). **

Conclusion

The findings suggest that a well-designed harness can act as a scalable and model-agnostic interface for embodied manipulation.

Improvements for AI systems

Based on the scientific paper Guava: An Effective and Universal Harness for Embodied Manipulation, here are specific improvements that can be made to AI systems, leveraging the insights from this research:


  1. Dominance of a Modular Harness Architecture (Guava Framework)

  2. Iterative Reasoning-Action Loops (ReAct Style)

  3. Semantic Action Abstractions for High-Level Planning

  4. Multimodal Observation Integration (Vision + Text State Representations)

  5. Data-Efficient Knowledge Distillation into Compact Models

Specific Improvements and Capabilities of the Improved AI System:

  1. Dominance of a Modular Harness Architecture (Guava Framework):

  2. Iterative Reasoning-Action Loops (ReAct Style):

  3. Semantic Action Abstractions for High-Level Planning:

  4. Multimodal Observation Integration (Vision + Text State Representations):

  5. Data-Efficient Knowledge Distillation into Compact Models:

The improved AI system, powered by the Guava harness, can perform the following specific capabilities:

  1. Perform complex, long-horizon manipulation tasks in real-world environments with high success rates (comparable to frontier proprietary models) even when trained on significantly less data (fewer than 2K trajectories).

  2. Demonstrate robust recovery from execution failures (e.g., failed grasps, object shifts, control issues) by continuously interleaving perception, reasoning, and action execution in a closed-loop process.

  3. Achieve strong generalization to unseen objects and novel instructions (Out-of-Distribution tasks) by relying on semantic task decomposition rather than low-level geometric reasoning.

  4. Operate efficiently using compact open-source models (e.g., 4B parameters) while maintaining high performance, overcoming the computational cost of deploying massive frontier models for every inference step.

  5. Exhibit strong zero-shot transfer capabilities from simulation to the real world, effectively separating high-level semantic planning from low-level perception and control to mitigate sim-to-real gaps without extensive real-world fine-tuning.

  6. Inference can be made more token efficient compared to large proprietary models (e.g., GPT-5.4) while maintaining competitive performance, leading to lower operational costs for agentic manipulation systems.

Sources

Related papers