SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis
cs.AI, cs.RO
Submitted: 2025-03-13
Updated: 2026-09-16
Journal ref: IEEE Robotics and Automation Letters, 2026, pp. 1-8
Code: https://github.com/jinlab-imvr/SurgRAW
License: http://creativecommons.org/licenses/by/4.0/
The gist: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding.
Terminology
Abstract
Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding. Most existing surgical AI methods rely on isolated, task-specific models, leading to fragmented pipelines with limited interpretability and no unified understanding of RAS scene. Vision-Language Models (VLMs) offer strong zero-shot reasoning, but struggle with hallucinations, domain gaps and weak task-interdependency modeling. To address the lack of unified data for RAS scene understanding, we introduce SurgCoTBench, the first reasoning-focused benchmark in RAS, covering 14256 QA pairs with frame-level annotations across five major surgical tasks. Building on SurgCoTBench, we propose SurgRAW, a clinically aligned Chain-of-Thought (CoT) driven agentic workflow for zero-shot multi-task reasoning in surgery. SurgRAW employs a hierarchical reasoning workflow where an orchestrator divides surgical scene understanding into two reasoning streams and directs specialized agents to generate task-level reasoning, while higher-level agents capture workflow interdependencies or ground output clinically. Specifically, we propose a panel discussion mechanism to ensure task-specific agents collaborate synergistically and leverage on task interdependencies. Similarly, we incorporate a retrieval-augmented generation module to enrich agents with surgical knowledge and alleviate domain gaps in general VLMs. We design task-specific CoT prompts grounded in surgical domain to ensure clinically aligned reasoning, reduce hallucinations and enhance interpretability. Extensive experiments show that SurgRAW surpasses mainstream VLMs and agentic systems and outperforms a supervised model by 14.61% accuracy. Dataset and code is available at https://github.com/jinlab-imvr/SurgRAW.git.
Sources
- SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
- CARES: Collaborative Agentic Reasoning for Error Detection in Surgery
- GPT-4 Technical Report
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
- Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
- MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow
- MedGemma Technical Report
- EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
- SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model
- SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
- Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
- General surgery vision transformer: A video pre-trained foundation model for general surgery
- MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Qwen3 Technical Report
- LLaVA-OneVision: Easy Visual Task Transfer
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection