RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance

arXiv:2609.39384 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance".

Dev: RoboAssist presents an agent-based framework for interactive human–humanoid planning designed to enable long-horizon surgical assistance by coordinating robot actions with evolving human activities while ensuring safety across planning and execution.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Well, what makes this paper stand out from other work in this space is its focus on interactive planning, which means the robot isn't just following a fixed script but actively reasoning about what the surgeon needs next.

Dev: I agree; that kind of dynamic interaction pushes the limits on loop rates and latency because you can't just run a standard planning cycle and expect perfect real-time alignment with human input.

Taro: From my side, I wonder if this agent-based structure actually gives it the necessary flexibility to handle unexpected environmental surprises during surgery, not just planned task changes.

The paper's summary: Rosa: The core concept they are presenting in "RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance" is this asymmetric dual-track representation that neatly separates what the human is doing from what the robot needs to do.

Dev: That separation sounds smart, especially because it allows them to treat the human process states as something partially observed while treating the robot tasks as executable sequences.

Taro: But how does this dual-track system actually maintain a meaningful connection between those two tracks, especially when dependencies are constantly shifting during a long operation?

The paper's improvements: Rosa: They suggest that the main improvement is this mechanism where they link those two tracks using "evidence-gated dependencies," which means a task only proceeds if the necessary human progress evidence is confirmed.

Dev: That sounds like a great way to manage uncertainty, because it stops the robot from blindly executing tasks based on incomplete information about the surgeon's current status.

Taro: I think that gating mechanism is key when things misbehave; it suggests that if a requirement isn't met by the observed human state, the system knows exactly where to stop and re-evaluate.

Conclusion: Rosa: To wrap up this discussion on "RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance," we see a framework that really formalizes how an agent can handle the continuous feedback loop between human intent and robotic action.

Dev: It seems they've managed to keep the planning responsive by only replanning what's necessary when things actually change, which is crucial for keeping things running smoothly in real-time scenarios.

Taro: I think their focus on cross-layer safety architecture, which includes preventive navigation and fail-safe supervision, shows they aren't just thinking about the plan but also the physical execution risks involved in that surgery environment.

Jingwei Jia, Keyu Zhou, Jiewei Wang, Peisen Xu, Xingyuan Zhou, Liang Wang, Jiming Chen, Gaofeng Li

Hangzhou Dianzi University, China. · Zhejiang University, China. · New York University, USA. · Italian Institute of Technology (IIT), Italy

cs.RO

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: Project Web: https://roboassist.github.io

Project page: https://roboassist.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: RoboAssist presents an agent-based framework for interactive human–humanoid planning designed to enable long-horizon surgical assistance by coordinating robot actions with evolving human activities

Key concepts

Asymmetric Dual-Track Representation
This is a core design feature that separates the planning into two distinct paths: one for tracking what humans are doing (the Human Process Track) and another for tracking what the robot needs to do (the Robot Task Track). This separation prevents mixing unpredictable human actions with fixed robot actions, making coordination much cleaner.
Evidence-Gated Dependencies
This mechanism links the human track and the robot track. A dependency between a human step and a robot task is only considered active if specific evidence (like observation or confirmation) is present. This ensures that the robot doesn't plan for a human action until that action is actually confirmed or observed.
Residual Replanning
When new information invalidates an old requirement or dependency, the system doesn't restart everything. Instead, it identifies the point where things broke (the divergence point) and only regenerates the necessary future steps. This targeted approach keeps planning efficient and fast when changes occur.
Cross-Layer Safety Architecture
This is a three-level safety system protecting the robot at different stages: preventing dangerous movement near humans, reacting to unexpected human proximity, and having an independent supervisor monitor physical robot health in real-time. This layered approach ensures safety from planning through execution.

Terminology

Summary

RoboAssist presents an agent-based framework for interactive human–humanoid planning designed to enable long-horizon surgical assistance by coordinating robot actions with evolving human activities while ensuring safety across planning and execution. The core contribution is an asymmetric dual-track representation that separates partially observed human process states from executable robot task sequences, allowing the system to maintain online dependency alignment and perform targeted residual replanning when workflow requests change.

The gist: RoboAssist is an agent-based framework for interactive human–humanoid planning that integrates workflow reasoning, task coordination, and cross-layer safety.

Core Framework and Representation

RoboAssist is built around an asymmetric dual-track representation that separates partially observed human process states from executable robot tasks, linking them through evidence-gated dependencies. This design is crucial because it avoids mixing non-controllable process estimates with robot-controllable actions. The system maintains two distinct tracks: the Human Process Track (Ut) and the Robot Task Track (Rt).

The Human Process Track, denoted as Ut, defines human progress through nodes where each node includes a human state or action description, associated entities, and an assistance need, linked through observation-updated task memory (Mt). The Agent estimates procedural progress from observations and predicts subsequent actions. A condition is only confirmed when the criterion τi is satisfied by admissible observational evidence or explicit human confirmation; missing evidence leaves it unconfirmed.

The Robot Task Track, Rt, represents a sequence of executable robot task nodes, where each node includes the bound skill (kj), invocation parameters (θj), execution status (σj), and its precondition (pj) and expected effect (ej). The Agent generates this sequence from confirmed requirements and anticipated assistance needs using workflow knowledge.

Interactive Planning and Online Alignment

The framework achieves online task alignment by linking the two tracks through a Gating Map, Zt = (Ut, Rt, ϕt), where ϕt(j) indexes the human-process nodes on which task vj depends. This map provides traceable human–robot dependencies for consistency checking and residual replanning.

At each decision cycle, new observations update the state memory (Mt). The Agent then revises Ut+1 and checks the remaining robot tasks for changes in requirements or dependencies. For an unfinished task vj, gate satisfaction is evaluated by calculating gj (Ut+1) = i∈ϕt(j) τi(Ut+1). If a gate is unsatisfied, the task waits; if a physical precondition pj (s) is unmet, it is checked at dispatch.

Cross-Track Consistency and Residual Replanning

Residual replanning is triggered when new evidence invalidates a requirement or human-process dependency, or when a required precondition can no longer be established. The divergence point k∗ is defined as the earliest unfinished task whose requirement or human-process dependency is invalidated, or whose physical precondition can no longer be established by the retained prefix R<k∗t. The prefix R<k∗t is preserved, and only the affected suffix is regenerated to ensure new tasks remain compatible with the expected effects of preserved tasks.

Cross-Layer Safety Architecture

RoboAssist employs a three-layer safety design that coordinates protection across planning, handover, and runtime execution:

  1. Preventive surgeon-aware navigation regulation: This layer regulates locomotion by commanding vcmd t = (0, dt ≤ 0, α(dt) vnom t, dt > 0), where α(dt) is a clearance-dependent velocity scaling factor that approaches unity far from the protected region and issues a zero-velocity command at boundary contact or entry.

  2. Reactive interaction regulation for VLA handover: This layer separates task-oriented motion generation from interaction safety regulation. A VLM-based monitor evaluates the relative human–robot configuration, and if unexpected proximity is detected, it interrupts nominal behavior to issue an avoidance response.

  3. Fail-safe whole-body runtime supervision: An independent supervisor continuously monitors proprioceptive feedback using the anomaly flag et = eF t ∨ eV t ∨ eO t (force, joint velocity, and persistent oscillation detectors). If the latched emergency-stop state is activated (lt), it terminates the current skill, inhibits subsequent motion commands, and has the highest authority over other layers.

Experimental Evaluation

The framework was implemented on a Unitree G1 humanoid robot in long-horizon, multi-stage simulated surgical assistance scenarios involving multimodal interaction, instrument handling, medical material transport, navigation, and safe human–robot handover. The experiments compare RoboAssist against a Scripted FSM and a Reactive agent. In the request-order adaptation benchmark (where delivery order changes), RoboAssist achieved an 81.25% success rate compared to 0% for the scripted FSM, while significantly reducing mean replanning latency from 3.62 to 1.

Improvements for AI systems

Here are the specific improvements for AI systems based on RoboAssist, and what these improved systems can achieve:

  1. Improve long-horizon task coordination in complex human-robot environments (e.g., surgery). The system will transition from simple command execution to a sophisticated agent capable of reasoning about evolving surgical workflows, anticipating surgeon needs, and managing multi-stage procedures requiring sequential dependencies (like picking up an instrument, moving it to a handover table, and then handing it over).

  2. Implement robust online task adaptation via residual replanning. The improved system will maintain an asymmetric dual-track representation separating uncertain human process states from executable robot tasks. When human requests change or observations invalidate prior assumptions, the AI will only replan the affected suffix of the task sequence, drastically reducing planning latency and computational overhead compared to full-plan regeneration.

  3. Integrate a cross-layer safety architecture for surgical autonomy. The system will employ three distinct safety layers:

  4. Preventive navigation regulation (decelerating as it approaches a surgeon's protected operating region).

  5. Reactive interaction regulation during handovers (interrupting nominal motion if unexpected proximity or human movement is detected).

  6. Fail-safe whole-body runtime supervision (monitoring joint velocity, force feedback, and oscillations to trigger an emergency stop upon detecting anomalies).

These improvements result in an AI system capable of performing:

  1. Predictive instrument provisioning and material transport within surgical workspaces while maintaining a surgeon-aware safety margin.

  2. Seamless adaptation to dynamic workflow changes requested by the surgeon (e.g., changing the order of required deliveries) with near-instantaneous planning response times, ensuring task continuity without long stalls or complete re-planning cycles.

  3. Safe, reliable physical interaction during close-range handovers and high-precision manipulation by dynamically adjusting motion based on real-time human proximity and physical anomaly detection.

Sources

Related papers