OctoNest: Adaptive Cross-Device Execution through Stateful Control
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "OctoNest: Adaptive Cross-Device Execution through Stateful Control".
Tom: OctoNest proposes HRePlan, a hierarchical replanning framework for multi-device agents designed to handle failures robustly across heterogeneous environments.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We've seen how the paper introduces the OctoNest: Adaptive Cross-Device Execution through Stateful Control framework, and now we're looking at who put this together. The authors are a team of researchers from Shanghai Jiao Tong University, Southeast University, Tsinghua University, and others.
Jane: That’s quite an international collaboration; it shows that tackling these kinds of multi-device challenges requires expertise drawn from several different strong AI research institutions working together on the same problem space.
Lu: The diversity of the team is important because it suggests they're drawing on different specialized knowledge bases, which is necessary when you're dealing with the complexity of heterogeneous environments and various execution strategies like API or CLI.
Meng: I’m interested in how their specific backgrounds influenced the design choices; for instance, seeing how their focus areas translated into the separation between device-local strategy recovery and orchestrator-level global replanning is telling.
Lalam: I think what's interesting about this team is how they managed to unify those execution methods—API, CLI, and GUI—into a single framework, which speaks to the need for flexible interaction in real-world AI deployments.
Tom: So it’s not just one person with a great idea; it’s a coordinated effort to build this whole structure that bridges the gap between high-level planning and low-level device interaction.
Jane: Precisely; it demonstrates that solving these problems often requires combining deep expertise in planning theory with practical engineering knowledge of how different operating systems behave.
The paper's summary: Tom: Now let's look at the actual core of the OctoNest paper, which summarizes the framework itself. Essentially, it describes HRePlan as a hierarchical replanning framework that organizes cross-device execution into a closed-loop process involving a global plan, strategy execution agents, and the external environment.
Jane: So it sets up this loop where the Orchestrator keeps the big picture plan, while device-level Strategy Planners are responsible for coordinating API, CLI, and GUI agents to actually run things on the physical devices.
Lu: The key takeaway is that this hierarchy explicitly separates system-level recovery from device-level recovery; it allows the system to decide whether to fix a strategy locally or if it needs a complete global plan repair.
Meng: That separation means you can be very specific about what kind of failure is being handled at each level, which simplifies the overall design because you don't have one monolithic recovery mechanism trying to solve everything at once.
Lalam: The framework’s ability to distinguish between repairable local failures and those needing cross-device reassignment is central; it provides the necessary intelligence to make smart decisions about when to ask for help from the higher levels.
Tom: It seems like this intelligent decision-making process, driven by that failure abstraction, is what elevates it beyond simpler systems that just rely on coarse-grained retries or simple plan revisions.
Jane: Exactly; they use this abstraction to give the Orchestrator structured feedback about what's happening locally, making the global replanning much more informed and targeted when it does intervene.
The paper's improvements: Tom: Moving on, let’s focus on what specific improvements HRePlan offers over existing solutions. It introduces the Cross-Layer Failure Event, or CLFE, which is a compact abstraction designed to bridge local execution and global replanning efficiently.
Lu: The CLFE encapsulates five key elements: the identity of the failed subtask, the source device, a categorized failure type—strategy-specific, service-specific, or device-specific—and a summary of local strategy attempts with their observations for escalation.
Meng: That abstraction is powerful because it prevents the global context from being flooded with raw low-level traces; instead, it gives the Orchestrator just what it needs to decide on the next big move.
Lalam: This structured feedback mechanism allows for a more compact communication pathway, which I think is crucial for keeping the overall system performant and reducing token costs in long workflows.
Tom: Beyond CLFE, they also detail how the Strategy Planner operates in three states—CREATE, PROGRESS CHECK, and REPLAN—allowing it to select or revise execution strategies on the fly based on local observations.
Jane: That dynamic state management within the Strategy Planner is what allows it to react immediately when things aren't going as planned locally, enabling swift local strategy revision without needing global intervention.
Conclusion: Tom: So, to wrap up this discussion on OctoNest: Adaptive Cross-Device Execution through Stateful Control, the main implication is that we have a framework that moves beyond simple error handling to a sophisticated hierarchical recovery mechanism for multi-device agents.
Jane: It means agents can now effectively manage tasks spanning different devices by intelligently deciding whether to fix things locally or when they need a full reassignment or global plan repair based on the structured failure information provided by CLFE.
Lu: I think the future potential lies in how this capability allows AI systems to tackle incredibly complex, real-world scenarios that inherently involve coordinating actions across multiple physical machines without getting stuck in brittle recovery loops.
Meng: From my side, I see this as a significant step toward building more reliable agents for industrial applications where uptime and accurate task completion are paramount because they’ve shown how to model the state space of heterogeneous failures effectively.
Lalam: For me, the most exciting aspect is that this paper suggests we can build AI systems capable of handling those intricate, multi-step workflows across Linux and Android environments with a much higher degree of reliability than before.
Tom: It’s clear that OctoNest offers a structured way to manage the inherent unpredictability of real-world, multi-device interactions in an organized manner. We're excited to see how this concept evolves into even more capable agent systems.
School of Artificial Intelligence, Shanghai Jiao Tong University
cs.CL
Submitted: 2026-06-18
Updated: 2026-09-30
Code: https://github.com/openclaw/openclaw
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 88/100
The gist: OctoNest proposes HRePlan, a hierarchical replanning framework for multi-device agents designed to handle failures robustly across heterogeneous environments.
Key concepts
- Hierarchical Replanning Overview
- HRePlan structures cross-device execution as a closed loop involving a global plan, strategy agents, and the environment. The Orchestrator manages the overall plan, while device-level Strategy Planners handle local execution. This separation ensures system-level recovery is distinct from device-specific fixes.
- Orchestrator Functions
- The Orchestrator acts as a system planner with three modes: Ocreate (creating sequential subtasks), Oappend (appending context based on success), and Oreplan (global replanning upon failure). It decides when to intervene globally versus letting local strategies handle issues.
- Cross-Layer Failure Event (CLFE)
- CLFE is a compact abstraction that bridges local execution and global replanning. Instead of sending full low-level traces, it summarizes failures into five key elements: subtask identity, source device, failure type, and a summary of local attempts. This allows the Orchestrator to make informed decisions without being overwhelmed by detail.
- Strategy Planner Functions
- The Strategy Planner is the device-level planner that selects and revises execution strategies (API, CLI, or GUI). It operates in CREATE, PROGRESS CHECK, and REPLAN states. It decides whether to execute a step locally or escalate the problem if a viable local path remains.
Terminology
Summary
OctoNest proposes HRePlan, a hierarchical replanning framework for multi-device agents designed to handle failures robustly across heterogeneous environments. This framework addresses the limitation of existing systems that rely on coarse-grained recovery by separating device-local strategy recovery from orchestrator-level global replanning. By introducing a platform-independent strategy abstraction and a compact cross-layer failure abstraction (CLFE), HRePlan enables agents to distinguish between failures that can be repaired within the current device and those requiring cross-device replanning, leading to substantially higher task completion rates and reduced token costs in complex workflows.
Hierarchical Replanning Overview
HRePlan organizes cross-device execution as a closed-loop process involving a global plan, strategy execution agents, and the external environment. The Orchestrator maintains the global plan, while device-level Strategy Planners coordinate API, CLI, and GUI execution agents against heterogeneous environments. This hierarchy separates system-level recovery from device-level recovery. The Strategy Planner handles failures that can be resolved within the current device by changing or continuing local strategies, while the Orchestrator intervenes only when device-level feedback indicates that the remaining work requires cross-device reassignment, downstream context revision, or global plan repair.
Orchestrator Functions
The Orchestrator is the system-level planner in HRePlan and operates in three modes:
-
Task Creation (Ocreate): Decomposes the instruction and device profiles into an ordered sequential subtask chain, specifying a concrete instruction and a target device selected from available devices.
-
Information Append (Oappend): After a subtask succeeds, it examines the result to decide whether later subtasks require additional context, appending information only to downstream subtasks that depend on the completed result.
-
Global Replanning (Oreplan): When a subtask fails, it receives structured failure evidence and rewrites the remaining chain by retrying with a modified instruction, rerouting to another device, inserting recovery subtasks, preserving reusable partial progress, or declaring global failure when no viable recovery path remains.
Strategy Planner Functions
The Strategy Planner is the device-level planner responsible for selecting and revising execution strategies within a single device. It maps the assigned subtask and previous observations to one of three decisions: EXECUTE(π, x), DONE(yj), or ESCALATE(cj). The available strategies are API (for structured functions), CLI (for local shell operations), and GUI (when structured strategies are unavailable or unsuitable). The planner operates in three states: CREATE, PROGRESS CHECK, and REPLAN. In the CREATE state, it selects an execution strategy; in the PROGRESS CHECK state, it examines results to decide if another strategy step is needed; and in the REPLAN state, it reacts to a failed local execution step by selecting another available strategy when a viable local path remains.
Cross-Layer Failure Event (CLFE)
HRePlan introduces the Cross-Layer Failure Event (CLFE), denoted as cj, as a compact, planning-oriented failure abstraction. CLFE bridges local execution and global replanning by exposing recovery-relevant facts to the Orchestrator without overloading the global context with full low-level traces. CLFE encapsulates five key elements:
-
The identity of the failed subtask.
-
The source device.
-
The categorized failure type (strategy-specific, service-specific, device-specific).
-
A summary of local strategy attempts with their observations and reasoning for escalation beyond the current device context.
Execution Agents
HRePlan implements each execution strategy with a corresponding strategy execution agent: Ad,π(x) → (status, y, ω), where x is the local execution instruction, status indicates completion or failure, y is the returned result, and ω is local execution evidence such as observations or tool traces. To provide comprehensive coverage across heterogeneous environments, HRePlan instantiates three complementary agents: an API Agent for reliable service access; a CLI Agent for local computation and file-system manipulation; and a GUI Agent for broad, user-interface interaction when structured access is insufficient.
Evaluation Benchmark
To evaluate hierarchical recovery, HRePlan is tested on HeraBench, a fault-injected benchmark that constructs cross-device workflows over Linux and Android devices. HeraBench injects strategy- and device-level failures that require both local strategy recovery and cross-device replanning. The evaluation metrics include Task Completion (percentage of required external-state postconditions satisfied), Instruction Adherence (ratio of dataset gold steps successfully fulfilled), and Perfect Pass, which requires achieving both 100% completion and 100% adherence. Efficiency is measured by average total tokens consumed per episode (Tok./Ep) and Expected Cost per Perfect Pass (Tok./PP).
Improvements for AI systems
Here are the specific improvements and capabilities derived from the H-RePlan framework described in the paper:
The proposed system, based on H-RePlan, improves AI systems by shifting from coarse-grained task management to a fine-grained, scope-aware hierarchical recovery mechanism for cross-device agent workflows.
Here are the specific improvements and what the improved AI system can do:
These improvements enable the system to perform: 174 variants of complex, multi-step tasks across heterogeneous environments (Linux and Android), such as submitting an expense claim involving mobile app data collection, local file organization on a laptop, and cloud uploads—all while maintaining high fidelity to the original user intent.
Sources
- MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- CoAct-1: Computer-using Multi-Agent System with Coding Actions
- Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
- MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
- API Agents vs. GUI Agents: Divergence and Convergence
- UFO3: Weaving the Digital Agent Galaxy
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering