Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

arXiv:2608.16889 · cs.RO, cs.AI, cs.CV · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Don't Drop the BATON".

Dev: Long-horizon robot manipulation often fails because errors compound across many contact-rich skills,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, if I’m following up on what we just touched on, BATON is essentially proposing a way to break down long tasks into smaller pieces so the AI can explore them one by one instead of trying to solve the entire thing end-to-end at once. It’s about making every small step an independent unit of exploration.

Dev: That decomposition changes the cost calculation significantly, because they argue that this makes the exploration cost additive, scaling with something like T times K, where T is episodes per stage and K is the number of stages, instead of multiplying it by the whole task length.

Taro: From an autonomy perspective, I see how this helps when things go wrong; if one subtask fails because of a state issue or a bad grasp, you know exactly which unit caused the problem, not some arbitrary point in the execution chain.

Rosa: That’s right; they are moving away from that vague "failure somewhere" feeling and towards something traceable, where every transition in that long sequence becomes an object that can be inspected and corrected.

Dev: The mechanism they introduce to manage these dependencies is what really sets it apart: the transition-aware memory which explicitly models how control moves between those subtasks across three specific types of transitions.

Taro: I’m paying close attention to those three types of transitions—the invocation, handoff, and lookahead—because that seems to be the key to managing those inter-subtask dependencies they mentioned.

Rosa: Right, because the paper says that simply having successful subtasks doesn't guarantee they will chain together correctly; the state left by one subtask often disturbs the next one’s required entry condition, which this memory tries to fix.

Dev: Specifically, they detail how the handoff transition records an "entry condition" to restore the state disturbed by the predecessor's residue when moving from one subtask to another, which is a crucial engineering detail for loop stability.

The paper's summary: Rosa: When we talk about improvements in "Don't Drop the BATON," the authors are really focusing on giving this agent a structured way to learn by making the subtask itself the primary object of exploration, which is a big conceptual shift.

Dev: Beyond just exploring subtasks, they’ve added that transition-aware memory system, which includes three specific contracts: invocation transition within a subtask, handoff across different ones, and lookahead across stages. These are explicit rules for how the AI should interact with its frozen VLA model.

Taro: The lookahead transition is particularly interesting to me; it means the agent can choose an execution strategy for the current step that considers what will be needed in future steps, which sounds like a smart way to handle long-term planning constraints.

Rosa: It allows the scheduler component of BATON to select a strategy that works for both where it is now and where it’s going, which addresses the issue of subtasks being interdependent in a way that traditional sequential learning can't see.

Dev: And they also detail how this hierarchical composition works, starting with decomposition into subtasks, followed by bootstrapping for new units lacking memory, and finally composing them outward "level by level" where related neighbors are chained first as trusted units.

Taro: I’m thinking about the practical implication of that composition—it means the system builds trust incrementally, only combining larger pieces once they have proven reliable at the smaller scale, which seems much safer than trying to learn one giant policy.

The paper's improvements: Rosa: So, to wrap up on "Don't Drop the BATON," it really boils down to treating long-horizon robot manipulation not just as a single skill chain, but as a series of verifiable subtasks connected by explicit transition contracts and hierarchical learning.

Dev: The main implication for us is that we can move towards systems where failure isn't just an uninformative crash but something diagnosable at the exact stage or seam where it broke, which is essential for building reliable control loops.

Taro: I think the real impact is in making these complex sequences tractable; if we can manage that cost additively instead of multiplicatively, it opens up possibilities for much longer and more intricate autonomous missions in unstructured environments.

Rosa: Absolutely, so by making every transition a first-class object and using this agentic subtask exploration method detailed in "Don't Drop the BATON," we get a framework that’s auditable and corrects itself at the right place.

Dev: We need to keep watching how they implement those handoff transitions under high-frequency execution; if latency creeps up during that state restoration, the whole additive cost benefit could disappear quickly.

Taro: It sounds like this framework gives us a much better handle on autonomy because it’s not just about executing the plan; it’s about understanding why the plan breaks at each handoff point.

Conclusion: Rosa: So we've been diving into "Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory," and to wrap up, this paper really shows how treating subtasks as units of exploration, combined with that transition-aware memory, lets the AI learn by writing all its adaptation into language memory.

Dev: It’s a neat way to keep the exploration cost additive rather than multiplying it by the whole task length, which is vital for keeping things stable during training runs.

Taro: I think this approach gives us a much clearer way to diagnose failures, because instead of some random point in a thousand-step trajectory failing, we can pinpoint exactly which stage or boundary caused the issue.

Rosa: Exactly. If one subtask fails due to a bad grasp, we know it’s that specific unit of exploration that needs refinement, not some part of the entire long sequence.

Dev: And those transition contracts—the invocation, handoff, and lookahead—they provide verifiable conditions for when the AI is allowed to move control between those stages.

Taro: That ability for the agent to select a strategy based on future requirements through that lookahead transition sounds like it really lets it handle situations where things go wrong in unexpected ways during execution.

Rosa: It’s impressive how this moves away from monolithic end-to-end models and gives us a structured way to build these complex sequences reliably.

Dev: The engineering aspect is that having those explicit handoff conditions means we can actually monitor the state residue between subtasks, which should help us debug latency issues in the loop rate.

Taro: For autonomy, this suggests we can tackle really intricate, multi-stage tasks that require a deep understanding of sequential dependencies without getting completely lost in the massive search space of a single long plan.

Rosa: It makes the whole process more auditable because everything is written into language memory, which is a big step toward creating more robust and correct robotic systems.

Dev: It certainly seems like it could translate well to deploying these on real hardware, provided those transition checks can be executed within the required loop rate constraints.

Taro: This work really sets a new standard for how we approach long-horizon planning in embodied agents; I'm curious to see how they apply these same principles to more dynamic or unpredictable environments next.

University of Southern California · University of Central Florida

cs.RO, cs.AI, cs.CV

Submitted: 2026-08-17

Updated: 2026-09-30

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 89/100

The gist: Long-horizon robot manipulation often fails because errors compound across many contact-rich skills, and current end-to-end models lack the ability to correct accumulated drift or understand how one

Key concepts

Agentic Subtask Exploration
This method breaks long tasks into smaller pieces so the AI explores them one by one instead of trying to solve the whole thing at once. Each small step becomes an independent unit of exploration.
Transition-aware Memory
This mechanism explicitly models how control moves between subtasks across three types of transitions: invocation, handoff, and lookahead. It records entry conditions to restore the state disturbed by a preceding subtask.
Additive Cost Calculation
By decomposing tasks into subtasks, the exploration cost becomes additive, scaling with T times K (episodes per stage times number of stages). This is better than multiplying the cost by the whole task length.
Lookahead Transition
This transition allows the agent to choose an execution strategy for a current step that considers what will be needed in future steps. It helps the scheduler select a strategy that works for both where it is now and where it’s going.

Terminology

Summary

Long-horizon robot manipulation often fails because errors compound across many contact-rich skills, and current end-to-end models lack the ability to correct accumulated drift or understand how one subtask constrains the next. This paper presents BATON, a novel coding agent framework that addresses these failures by treating the subtask as the unit of exploration and equipping it with transition-aware memory to manage inter-subtask dependencies. By decomposing long-horizon tasks into short-horizon, verifiable steps and explicitly modeling how control passes between stages, BATON achieves superior performance in complex manipulation benchmarks by making every transition a first-class object.

Agentic Subtask Exploration

BATON addresses the multiplicative cost of whole-task exploration by making the subtask the unit of exploration. Instead of exploring a thousand-step chain end-to-end, BATON decomposes the task into a sequence of short-horizon subtasks and explores each one in an inexpensive regime where its solution is stored in memory. This approach changes how failure is attributed: A long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost thus becomes additive (T · K), and every failure is attributed to a single stage. This allows the system to learn by test time exploration with no parameter updates, ensuring that the cost scales additively in the number of new subtasks and boundaries rather than multiplicatively.

Transition-Aware Memory

To solve the problem where neighboring subtasks are coupled through execution, BATON equips its exploration with a transition-aware memory that explicitly models how control moves between stages. This memory consists of three kinds of transitions:

  1. The invocation transition (within a subtask): The agent and the frozen VLA interact only after a wrist view confirms that the scene is ready.

  2. The handoff transition (across different subtasks): This records an entry condition to restore the state disturbed by the predecessor's residue.

  3. The lookahead transition (across different subtasks): This determines the execution strategy whose outcome the successor can inherit, allowing a scheduler agent to select a strategy that works for both current and upcoming subtasks.

Hierarchical Subtask Composition

BATON structures its learning through hierarchical composition, moving from individual verified units to the full task. The process involves:

  1. Decomposition: Drafting a decomposition of the long-horizon instruction into subtasks and boundaries, consulting memory to inherit previously solved units.

  2. Bootstrapping: For each new subtask lacking memory, the Explorer runs a bootstrapping loop—trialing staging orders, pre-contact poses, and VLA ACT invocations under the wrist-camera gate.

  3. Composition: Solved subtasks are composed outward, level by level, where related neighbors are chained first as trusted units before combining larger compositions.

Core Contributions and Results

The core contributions of BATON include:

** Presenting BATON, the first coding agent to carry long-horizon robot manipulation end to end over a frozen VLA, learning by test-time exploration with no parameter updates.**

(1) Hierarchical subtask exploration for, which keeps exploration cost additive.

(2) Transition-aware memory for, which attaches checkable conditions to the invocation, handoff, and lookahead transitions.

On the RoboMemArena benchmark of thousand-step household tasks, BATON improved task success rate by 11.6% and cumulative success rate by 14.9% over the current State-of-the-Art (SoTA). The results demonstrate that this approach allows failures to be localized: A failure names the stage that caused it instead of voiding a thousand-step episode. This contrasts with previous methods where errors compounded, leading to policy drift and uninformative failures.BATON makes every transition of a long-horizon task a first-class object.

Conclusion

BATON treats the problem as one of transitions rather than just skills or planners. It succeeds by making the subtask the unit of exploration and the transition the unit of memory, ensuring that exploration cost remains additive and that failures are diagnosable at a named stage or boundary. The system learns by writing all adaptation into language memory, allowing for a reusable, auditable, and corrected framework for long-horizon manipulation.

References

[1] Johan Bjorck et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025.

[2] Kevin Black et al. π0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024.

[3] Anthony Brohan et al.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems, based on the BATON framework described in the paper, and what these improved systems will be able to do:


The core improvement lies in transitioning from monolithic end-to-end models or simple memory augmentation to a structured, agentic architecture that treats long-horizon tasks as a sequence of verifiable subtasks connected by explicit, language-defined contracts.

Here are the specific improvements and capabilities:

  1. The system will be restructured around an LLM Coding Agent orchestrating a frozen Vision-Language-Action (VLA) model (like VLA ACT).

  2. This agent will replace monolithic planning by decomposing a long task into a sequence of short, manageable subtasks, exploring each one individually in the low-cost regime.

  3. The system will incorporate two explicit memory mechanisms:

  4. A hierarchical structure for exploration (to manage cost): Each subtask's success is stored as an abstract Skill Memory unit, making task failure traceable to a single stage rather than compounding across 1000+ steps.

  5. Transition-Aware Memory (the critical innovation): The system will explicitly record and enforce three types of cross-stage contracts:

  6. An internal Invocation Transition contract: It will only fire the VLA contact primitive when a specific, verifiable condition (e.g., wrist view confirmation) is met, preventing premature or erroneous contact.

  7. A Handoff Transition contract: It will record and restore the required entry state for the next subtask, actively correcting state residue left by the predecessor to ensure clean starts (e.g., ensuring a gripper is in a vertical position before picking up an object).

  8. A Lookahead Transition contract: It will allow the current subtask's execution strategy to be constrained by future requirements, enabling the agent to select the correct grasp or manipulation technique for the current step based on what is needed later (e.g., choosing a side grasp versus a top grasp because of a subsequent placement task).

The resulting improved AI system (BATON) can perform the following:

  1. Perform complex, multi-stage robot manipulations that are currently impossible for end-to-end VLA models to reliably chain together over long horizons (e.g., 1000+ steps).

  2. Achieve superior performance in tasks involving sequential dependencies where a single error compounds (e.g., Open drawer, pick up can, place can in drawer).

  3. Demonstrate cost-effective learning by only performing expensive contact-rich exploration on short subtasks, making the overall training process significantly faster and more budget-friendly than full end-to-end exploration.

  4. Exhibit high robustness against state drift and partial failures because its memory is structured around contracts rather than raw trajectories, allowing it to diagnose whether a failure was due to an inadequate skill or an incorrect state transition.

  5. Adapt dynamically to unforeseen constraints by using the lookahead transition to select the optimal execution strategy for any given subtask based on the known requirements of future steps in the plan.

Sources

Related papers