Towards the Harness of Embodied Agents

arXiv:2608.11246 · cs.AI, cs.LG, cs.RO · Submitted 2026-08-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Towards the Harness of Embodied Agents".

Jane: The paper was written by Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen et al. from Eastern Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everybody. I'm Tom, and we are kicking off a brand new paper today that has got me genuinely buzzing. We're looking at "Towards the Harness of Embodied Agents" from the team at the Eastern Institute of Technology in Ningbo.

Jane: And I'm Jane. Tom, I have to say, the moment I saw that title, I got excited, because "harness" is doing a lot of work there. It's not about building a better robot brain, it's about building the infrastructure around the brain.

Tom: Exactly. And that's a huge shift in thinking. For years, we've been obsessed with making the model itself smarter, but this paper is saying, wait, look at how coding agents actually succeed. They succeed because of the loop they live in, the tools they can grab, the feedback they get.

Jane: Right, and the authors are basically asking a really bold question. If that harness works so well for software, can we build the same kind of harness for a robot that has to move around and touch things in the real world?

Lu: You know, Jane, that question is exactly why I love this paper. I'm Lu, by the way, for our listeners. The authors take this very established paradigm from software and they ask what breaks when you try to port it over. And they find two fundamental gaps, not small bugs, but fundamental gaps.

Tom: And those two gaps are the heart of the whole thing. The first one is readability. In software, the agent can just read the code, grep it, diff it, it's all text. But in the physical world, the robot gets a thirty-frames-per-second stream of RGB-D images. It's just noise unless you structure it.

Jane: So the world doesn't come with a manual. You have to build the manual yourself. And the second gap is verifiability. When a coding agent runs a test, it gets an exit code, PASS or FAIL, and a stack trace if it breaks. A robot closes its gripper and it has no idea if it actually grabbed the cup or just closed on air.

Meng: And that's the part that keeps me up at night as an engineer. I'm Meng, by the way. In software, that feedback loop is free. The operating system gives it to you. But here, they have to build an entire evaluation module just to approximate what an exit code does in the physical world.

Tom: And that's what the paper calls "Evaluation as Exit Codes." It's such a clean way to put it. The robot needs a judge to tell it, "Hey, that grasp failed because you were too far away," not just "something went wrong."

Jane: So the title really is the thesis. It's not "Towards Better Robot Models," it's "Towards the Harness." The infrastructure is the product. And I think that's a really empowering way to think about robotics, because it means we can make progress without waiting for some magical superintelligent model to appear.

Lu: Precisely, Jane. And the math in the paper actually backs this up. They show that if each step in a task succeeds with eighty-five percent probability, a fifteen-step task under open-loop execution succeeds only about eight point seven percent of the time. But with a harness that has a good evaluator and allows retries, that number jumps to over ninety-nine percent.

Tom: Wait, let me just make sure I'm hearing that right. So without the harness, the robot is basically doomed to fail on any long task, but with the harness, it becomes nearly reliable, even though the underlying robot skills haven't improved at all?

Meng: That's the kicker, Tom. The harness is doing the heavy lifting. And that's why this paper is so important. It's saying, stop trying to make the perfect single-step robot, start building the system that can recover when that robot makes a mistake.

Jane: And that's the hook for our next segment. We're going to dig into the actual architecture they built, the agentic loop, the scene graph, and how they close that loop between the robot and the world. Stay with us.

Summary: Tom: Welcome back. We're deep in "Towards the Harness of Embodied Agents," and we've established that the harness is the star of the show. Now let's talk about what that harness actually looks like, because they didn't just theorize about it, they built it.

Jane: Right, and the architecture is called Thea. And the first thing that struck me, Tom, is that the core loop is almost identical to a coding agent. The model reads context, picks a tool, executes it, reads the result, and does it again. It's the same skeleton.

Lu: But the flesh on that skeleton is completely different. Take the context, for example. A coding agent reads files. Thea has to read the world. So they built something called the Scene Graph as Context. It's a persistent, symbolic map of everything the robot has seen, with stable IDs for objects.

Tom: So instead of the model trying to parse a blurry camera image every single turn, it gets a clean list. "Cup two is at this position, confidence one point zero, last seen two seconds ago." It's like giving the robot a database of the room.

Meng: And that's a massive engineering win. But I want to push back a little on the evaluation side, because that's where I was most skeptical. They have this evaluator that judges whether an action succeeded. How do they make that reliable enough to trust?

Jane: Well, they were smart about it. The evaluator is an independent module. It doesn't read the model's reasoning or its self-report. It only looks at the current camera observation and a per-tool post-condition, which is a description of what the world should look like if the action worked.

Lu: And that independence is crucial. The paper cites research showing that a model grading its own work is an unreliable judge. So they deliberately blind the evaluator to the model's intentions. It's like having a referee who didn't see the play call, only the result on the field.

Tom: And the verdict isn't just a boolean. It returns a structured failure reason. So if a grasp fails, it might say "the robot stopped too far from the bottle to grasp it." That's the stack trace for the physical world.

Meng: Okay, that's genuinely useful. But I have to ask about the numbers. How well does this evaluator actually perform? Because if it's wrong half the time, the whole loop falls apart.

Jane: They tested it on ninety trajectory checkpoints across three different robots. And it hit ninety-three point three percent accuracy. But here's the interesting part, the errors aren't random. It never misjudges a success as a failure. The errors are all false successes, where it thinks something worked when it actually failed.

Tom: And that's the dangerous direction, right? Because a false success means the robot moves on when it should retry. The paper is very honest about that being the key area for improvement.

Lu: But even with that imperfection, the math works. They plug in their measured accuracy of zero point nine three, a per-step success of zero point eight, and five retries, and a five-step task goes from a thirty-three percent success rate in open loop to ninety-one percent under the harness. That's the power of the loop.

Meng: So the evaluator doesn't have to be perfect, it just has to be good enough to make retries worthwhile. That's a really practical insight. You don't need a perfect judge, you need a judge that's better than random.

Jane: And that's the beautiful thing about this design. It's not waiting for perfect components. It's building a system that works with imperfect components and compensates for them. Which brings us to our next segment, where we talk about what actually emerges when you put all these pieces together.

Improvements: Tom: We're back with "Towards the Harness of Embodied Agents," and we've covered the architecture. Now let's talk about what happens when you actually let this thing loose in the real world, because the results are where it gets really fun.

Jane: And I love that they didn't just run the same boring pick-and-place task over and over. They let the agent loose on open-ended tasks and watched what emerged. And the first thing that emerged was long-horizon composition. The robot was navigating between rooms, manipulating objects, and doing it all in one continuous session.

Lu: But the most impressive part to me was the active perception. In one demo, the user asked for a power bank. The scene graph didn't have one. So the robot didn't give up. It navigated to a cabinet, opened the upper drawer, saw nothing, closed it, opened the lower drawer, and found the power bank there.

Meng: That's a big deal, Lu. That's not just following a script. That's the model deciding, based on the state of the world, that it needs to explore. It's using the scene graph to know what it doesn't know. That's a level of reasoning I didn't expect to see.

Tom: And it gets better. There was another demo where the robot failed to grasp an object. The evaluator came back with a reason, "you're too far away." And the robot didn't just retry the same way. It repositioned itself and tried again from a better angle. That's the recovery loop working exactly as designed.

Jane: And then there's the user interaction, which I think is so important. In one scene, the user asked for water, but there was no water in the scene graph. The robot didn't just grab something random. It asked the user, "No plain water, which drink should I bring?" And the user said orange juice, and the robot brought it.

Meng: That's the kind of behavior that makes a robot actually useful in a home. It's not just about executing commands, it's about knowing when to ask for clarification. That's a social skill, not just a technical one.

Lu: And I want to highlight the cross-embodiment portability, because that's what makes this a platform, not just a demo. They ran the same harness on three different robots. The Astribot S1, the AgileX Cobot Magic, and the Unitree G1 with dexterous hands. Nothing in the loop changed.

Tom: So the model, the loop, the evaluator, all of it stayed the same. They just swapped out the embodiment profile and the tool implementations. That's like taking the same operating system and running it on different hardware.

Jane: And that's the improvement this paper suggests. It's not a better robot, it's a better way to build robots. You can plug in a new policy as a new tool, and suddenly the whole system can do something new. Capability becomes additive.

Meng: And from an engineering standpoint, that's huge. It means you can have different teams working in parallel. One team improves the manipulation policy, another team improves the navigation, and the harness just integrates them. You don't have to retrain everything from scratch.

Lu: Exactly, Meng. And the paper even suggests that the loop itself could become a training target. Instead of just training a model to grasp, you train it to read a scene graph, choose a tool, and recover from failure. Those are learnable behaviors.

Tom: So the future isn't just a smarter model, it's a model that's been trained to live inside this harness. And that's a really exciting vision. But we've got to wrap up soon, so let's head to the conclusion and pull all of this together.

Conclusion: Tom: Alright, we've reached the end of our time with "Towards the Harness of Embodied Agents," and I have to say, this is one of those papers that changes how I think about the field.

Jane: It really does. And if I had to summarize the whole thing in one sentence, it's this. The reason coding agents work isn't because the models are magic, it's because they live in a well-built harness that gives them readable state and verifiable outcomes. And this paper shows you can build that same harness for robots.

Lu: And the two pillars of that harness are the Scene Graph as Context, which makes the physical world readable, and Evaluation as Exit Codes, which makes actions verifiable. Those are the two gaps the physical world doesn't fill for free, and Thea builds them from scratch.

Meng: And the results speak for themselves. The harness took a five-step task from a thirty-three percent success rate to ninety-one percent. And it did it without improving any of the underlying robot skills. That's the power of the loop, and that's what makes this paper so important.

Tom: And it's not just about the numbers. It's about the vision. The paper imagines a future where robots get more reliable over time, not because the model gets better, but because the harness learns. The memory system writes lessons into the tool descriptions, so the agent literally improves its own interface with use.

Jane: And that's the flywheel. Every task the robot completes teaches it something. Every failure it recovers from becomes a lesson. Over time, the harness gets smarter, the tools get better, and the robot becomes more capable. It's a self-improving system.

Lu: And that's the real legacy of this paper. It's not just a system, it's a paradigm. It's saying that embodied AI is entering the same shift that software agents already went through. The infrastructure matters as much as the model, if not more.

Meng: And as an engineer, that's incredibly empowering. It means we can build reliable robots with imperfect components. We don't have to wait for the perfect policy. We just need a good enough policy and a great harness around it.

Tom: Well said, Meng. And with that, we're going to say goodbye to "Towards the Harness of Embodied Agents." It's been a fantastic discussion, and I feel like we've only scratched the surface.

Jane: Absolutely. But that's the beauty of this field. There's always the next paper, the next idea, the next breakthrough. And we'll be here to talk about it. Thanks for listening, everybody. See you next time.

Tom: Take care, everyone.

Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen, Yizhong Ge, Jason Qin, Chengtai Li, Wentao Zhu

Eastern Institute of Technology

cs.AI, cs.LG, cs.RO

Submitted: 2026-08-03

Updated: 2026-08-13

Comments: Project page: https://eit-hai.github.io/thea

Code: https://github.com/EIT-HAI/Thea

Project page: https://eit-hai.github.io/thea

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 70/100

The gist: The paper presents Thea, a harness for embodied agents that applies the paradigm of coding agents to the physical world.

Key concepts

Agentic Loop
The core structure of the system, which is identical to software coding agents. It involves a continuous cycle where the model reads context, selects a tool, executes an action in the environment, and then processes that result to enable repeated operation.
Scene Graph as Context
A persistent symbolic map of the robot's environment. Instead of parsing raw sensory data like camera images, this provides a structured list of objects with stable IDs and locations, giving the robot a readable database of its surroundings.
Evaluation as Exit Codes
The method used to determine if an action succeeded. An independent module judges the outcome by comparing the current camera observation against a defined post-condition, providing structured failure reasons similar to how software provides a stack trace.

Terminology

Summary

The paper presents Thea, a harness for embodied agents that applies the paradigm of coding agents to the physical world. The core argument is that what an agent achieves depends not on the model alone, but on the infrastructure around it, and the paper asks whether the same paradigm extends to embodied agents in the physical world.

The paper identifies two fundamental gaps that distinguish the physical world from software environments. The first is Readability of the World: "A software environment is a designed artifact, built by humans for machines to read and execute... The physical world is not a designed artifact. It simply exists. An embodied agent perceives by sensing. Its input is a 30 fps stream of RGB-D frames, continuous, high-dimensional, and unstructured. The second is Verifiability of Outcomes": "Software environments possess an elegant property... every action has a natural termination signal and an explicit success/failure judgment... The physical world offers none of this. A Vision-Language-Action (VLA) policy outputs a continuous action stream with no built-in 'done' signal. There is no exit code; the world does not report 'grasp succeeded.'"

To bridge these gaps, Thea introduces two key components. "Scene Graph as Context restores readability: a persistent, structured model of the scene that the language model can read and reason about as a coding agent reads source code. Evaluation as Exit Codes restores verifiability: it detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause, closing the loop that the physical world otherwise leaves open."

The system architecture inherits the core components of coding agents, modified for the physical world. The agentic loop is the same loop that drives coding agents: the same accumulating messages, the same termination convention. Context engineering assigns each model-visible input both a representational role and a lifetime, so stable operating knowledge, current physical evidence, and historical trace occupy distinct parts of the context window: resident, refreshed, and accumulated. The tool protocol defines one interface and one result envelope for every tool, where each tool file carries the tool's full contract: the inputSchema says how to call it, the description says when, and the post-condition says how its outcome will be judged. Skills are a directory holding a SKILL.md file that injects domain knowledge. Memory is a lifecycle rather than a single store, with Task Notes for the active task, MEMORY.md for cross-session knowledge, and tool experience/ files for tool-specific lessons. Safety is built into the machinery rather than into the model's behavior, with deterministic checks in hooks and a safety filter on base motion.

The paper provides a formal reliability analysis. For an n-step task with per-step success probability p i, open-loop success is P open = ∏ p i. Under the harness with an evaluator of accuracy α and up to k retries per step, the task reliability becomes P harness = ∏ p i·α·(1-r i k)/(1-r i), where r i is the retry probability. The formula shows that reliable evaluation (α → 1) is the primary design challenge, and each additional retry yields geometrically less improvement, and the limit is directly capped by α.

Experiments were conducted on three real robot platforms: Astribot S1, AgileX Cobot Magic, and Unitree G1. Tasks were organized into three levels: L1 (short-horizon manipulation), L2 (long-horizon manipulation), and L3 (navigation and manipulation). Thea was compared against end-to-end policies (ACT, LingBot-VLA-V2, π0.5), a Coding-as-Policy baseline (CaP-X), and a hierarchical planning method (SayCan). The paper reports that Thea achieves the highest success rate at all three levels, with the gap widening from L1 to L3 as task complexity increases.

The evaluator accuracy was measured on 90 trajectory checkpoints across the three embodiments, achieving an average accuracy of 93.3%. The paper notes that the remaining errors all run in one direction: 6.7% of failed and 13.3% of in-progress checkpoints are judged successful, while no successful checkpoint is misjudged, identifying this as the harmful kind of error.

The paper demonstrates several emergent capabilities: long-horizon task composition, active perception (searching cabinet drawers for occluded objects), failure recovery (using evaluator failure reasons to reposition and retry), user interaction (querying the user when the requested object is unavailable), and cross-embodiment portability (the same harness running on all three robots with embodiment-specific profiles).

The paper concludes that embodied AI is entering the paradigm shift that software agents have already undergone, where the layer between the model and its environment has come to matter as much as the model itself. It argues that an embodied agent needs such a system, rather than just a capable model, even more than a coding agent does, as what it faces is far more complex. The paper envisions that reliability becomes a systems problem, capability becomes additive, and the loop becomes a training target for future embodied foundation models.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


  • What to build: A persistent, object-centric scene graph that maintains symbolic state (object refs, coarse 3D positions, confidence, freshness, container contents, robot pose/holding state) across turns.

  • How: Integrate a perception backend (RGB-D, LiDAR) that updates the graph; gate updates by an evaluator's confirmed outcomes; render a compact scene graph brief into the model's context each turn; provide ref-keyed query tools (get object relations, get image).

  • Result: The AI can reference the world symbolically (e.g., navigate to("cup 2")), resolve vague user requests (the red cup) by querying stored images, and maintain a coherent global picture without reconstructing state from raw frames.

  • What to build: An independent evaluator module that, after every manipulation, judges success/failure/in-progress using only (a) the current observation and (b) a per-tool post-condition. It returns a three-state verdict plus a structured failure reason.

  • How: Trigger via a post-execution hook (not model-discretion); use a separate VLM (e.g., Qwen3.7-Plus) that never sees the model's reasoning; feed verdicts into scene-graph updates and memory consolidation.

  • Result: The AI can detect failed grasps, reposition, and retry; it can distinguish too far to reach vs. object slipped vs. occluded and choose the correct recovery. This closes the loop that the physical world otherwise leaves open.

  • What to build: A reactive loop (one tool call per turn, observe, decide) where every capability (navigation, manipulation, perception, user interaction) is a callable tool with a contract: name, description, inputSchema, and a hidden POST CONDITION.

  • How: Register tools via a one-line decorator; wrap policy backends (ACT, π0.5, CaP-X) as subprocess MCP tools; enforce pre/post hooks (safety preflight, auto-evaluation).

  • Result: The AI can compose long-horizon tasks (navigate → adjust pose → grasp → place) at runtime, adapt to unexpected states, and recover from failures without hand-coded recovery routines.

  • What to build: A context builder that separates resident (system prompt, memory, embodiment profile, tool definitions), refreshed (scene graph brief, current observations), and accumulated (instructions, task notes, model responses, tool results) content.

  • How: Regenerate refreshed context before each decision; append task notes after each tool result; compact only the accumulated lifetime when nearing the context window.

  • Result: The AI always reads current physical evidence, never relies on stale observations, and maintains a concise working record of its own progress.

  • What to build: A three-tier memory: (a) Task Notes (per-task notepad), (b) MEMORY.md (cross-session preferences/conventions/lessons), (c) tool experience/*.md (per-tool success/failure lessons).

  • How: Consolidate at task end via an LLM call that splits entries; append tool-experience summaries to each tool's description at task start; read memory into resident context.

  • Result: The AI improves its own tool interface over time—it learns where a grasp policy actually succeeds, and applies that knowledge to future calls without retraining.

  • What to build: A replaceable document describing the active body: operational envelope (footprint, mobility, reachable workspace), perception configuration (sensor modalities, model-visible views), and base-relative positions (cameras, grippers).

  • How: Load the profile at deployment; keep it in resident context; swap profiles to port across robots (Astribot S1, AgileX Cobot Magic, Unitree G1).

  • Result: The same harness drives different robots; the AI knows its reach limits, which cameras to query, and how to interpret spatial measurements correctly.

  • What to build: Deterministic pre-execution hooks that (a) read fresh clearance measurements before any base motion, (b) block navigation when no direction is admissible, (c) clamp translations to safe distances.

  • How: Place in the tool-execution pipeline, not in the model's discretion; return failures through the normal result envelope.

  • Result: The AI cannot be persuaded to move into a wall or collide with an obstacle, even if the model's reasoning is flawed.

  • What to build: query user (blocking clarification) and notify user (non-blocking updates) as first-class tools.

  • How: Register them like any other tool; let the model decide when to call them; attach scene-graph refs and views to questions so answers ground subsequent actions.

  • Result: The AI asks for clarification when a request is ambiguous (e.g., no plain water—which drink?), and keeps the user informed without stalling.

  1. Complete long-horizon physical tasks (e.g., find a power bank and place it on the desk) by navigating, searching drawers, grasping, and delivering—with recovery from failures—at a success rate of 91% on 5-step tasks (vs. 33% open-loop), per the paper's model.

  2. Actively explore when a target is absent from the scene graph: navigate to a cabinet, open drawers in sequence, update the graph with newly visible objects, and then proceed.

  3. Recover from manipulation failures using evaluator-provided reasons: if a grasp fails because the robot is too far, it repositions and retries; if an object slips, it re-grasps with a different approach.

  4. Resolve ambiguous user requests by querying stored images and relations (e.g., the red cup → get image on candidate refs) or by asking the user directly when evidence is insufficient.

  5. Port across different robot bodies without redesign—swap the embodiment profile and tool implementations, and the same loop, context, and evaluator logic works.

  6. Improve its own tool descriptions over time: after each task, it writes lessons (e.g., lower drawer opens better when the base is 5 cm closer) into tool-specific memory, which is then injected into the tool's description on the next task.

  7. Operate safely in cluttered environments: it plans with clearance measurements in context, and the harness blocks any motion that would collide, regardless of the model's intent.

  8. Keep the user in the loop: it notifies before long actions, asks for permission before touching personal items, and asks for alternatives when a requested item is unavailable—all as ordinary tool calls.

These improvements are directly implementable from the paper's architecture (Section 3) and validated by its experiments (Section 5), which show the harness outperforming end-to-end policies, coding-as-policy baselines, and hierarchical planning on real robots across tasks of increasing complexity.

Abstract

The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We ask whether the same paradigm extends to embodied agents in the physical world. We present Thea, a harness in which an agentic loop orchestrates robot capabilities, each wrapped as a callable tool. It inherits the core components of coding agents, modified as the physical world requires. The world, however, withholds two abilities that software grants for free: reading the state of the world, and judging the outcome of an action. To bridge these gaps, Thea introduces Scene Graph as Context, a persistent, symbolic representation of the world, and Evaluation as Exit Codes, which detects when an action should terminate, judges whether it succeeded, and on failure diagnoses the cause. Together they close the loop between the agent and the physical world. Rich behaviors then emerge from the composition of tools, and the closed loop carries long-horizon tasks to completion in real environments.

Sources

Related papers