MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery

arXiv:2602.12407 · cs.RO, cs.CV, cs.LG · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery".

Jane: The paper was written by Keshara Weerasinghe, Seyed Hamid Reza Roodabeh, Andrew Hawkins, Zhaomeng Zhang, Zachary Schrader et al. from University of Virginia and University of Virginia Health System.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and as always, Jane is here with me. Today we're looking at a paper that's going to make a lot of robotics labs very happy. It's called "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery."

Jane: And Tom, I have to say, the first thing that struck me is the team behind this. It's from the University of Virginia, and they've got engineers working side by side with a thoracic and cardiovascular surgeon. That mix is exactly what you need when you're trying to solve a problem in the operating room.

Tom: Absolutely. And the problem they're tackling is one that anyone who's tried to do research on surgical robots knows all too well. The robots themselves, like the da Vinci system, are these closed, proprietary boxes. You can't just plug in and get the data you need.

Jane: Right. If you want to study how a surgeon moves, or how they press the pedals, or what the robot's arms are doing, you're usually stuck. The manufacturer doesn't give you easy access to that internal telemetry. So this team built their own system to capture it from the outside, without touching the robot's software at all.

Tom: And that's the "MiDAS" part. It stands for Multimodal Data Acquisition System. They're using electromagnetic trackers on the surgeon's fingers, a depth camera watching their hands, force sensors on the foot pedals, and the surgical video feed. All of it synced together in real time.

Jane: What I love is that they validated it on two very different robots. They tested it on the Raven-II, which is this open-source research robot, and then they took it to a clinical da Vinci Xi system during a real surgical training bootcamp. That's a huge deal because it proves the system isn't tied to one platform.

Tom: And they're giving it all away. The code, the hardware designs, and the datasets they collected, including a brand new one of surgeons doing hernia repair on realistic tissue models. That's a gift to the research community.

Jane: It really is. And it means that labs that don't have access to expensive, proprietary robots can still do meaningful research on surgical skill and safety. They just need a MiDAS setup and a robot to practice on.

Tom: So stick around, because in the next segment we're going to dig into what they actually collected and why that hernia repair dataset is such a big deal.

Summary: Tom: Welcome back. We're still talking about "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." And Jane, I want to get into the actual data they collected, because there's a lot here.

Jane: There really is. They ran two big experiments. On the Raven-II, they had people do the classic peg transfer task, moving little rings from one peg to another. That gave them fifteen trials and about thirty-six minutes of data, all annotated with seven different gestures like "approach peg" and "transfer peg."

Tom: And that's the dry-lab stuff. But then they did something much more ambitious. They went to a surgical bootcamp at the University of Virginia hospital and recorded surgeons doing actual hernia repair suturing on these high-fidelity tissue models made by a company called KindHeart.

Jane: Those models are made from real porcine tissue, so they feel much closer to real surgery than a plastic trainer. They got seventeen trials, over two hundred minutes of annotated data, with eight different suturing gestures. That's a lot of careful labeling work.

Tom: And the gesture taxonomy itself is interesting. They didn't just make it up. They built it with an expert surgeon, and it captures things like "orient needle," "push needle through tissue," and "make a C loop." These are the actual steps a surgeon goes through when closing a hernia.

Jane: What I find really impressive is the validation they did. They compared their external sensors against the robot's own internal kinematics on the Raven-II. The electromagnetic hand trackers matched the robot's motion really well, with cosine similarity above zero point eight on the position data.

Tom: And the foot pedal sensors, they compared those against the ground truth too. They got an F1 score of zero point eight five on the Raven-II and zero point seven eight on the da Vinci. So the system is actually capturing what's happening, not just guessing.

Jane: Exactly. And that's crucial because it means researchers can trust this data. They can use it to train models for gesture recognition, for skill assessment, even for detecting errors during surgery. And they don't need to hack into the robot to do it.

Tom: Which brings us to the next question. How well do these external signals actually work for recognizing what the surgeon is doing? That's what we're going to dig into next.

Improvements and Experiments: Tom: We're back with "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." And Jane, we just talked about the data collection. Now I want to talk about what they did with it, because they ran some serious experiments.

Jane: They did. They took their external hand tracking data and fed it into two state-of-the-art gesture recognition models. One is called MTRSAP, which is a transformer-based model, and the other is MS-TCN++, which is a temporal convolutional network.

Tom: And the results are pretty striking. On the peg transfer task, the electromagnetic hand tracking data performed almost as well as the robot's own internal kinematics. We're talking about an F1 score of zero point eight six versus zero point eight seven for the robot data. That's a tiny gap.

Jane: That's a huge finding. It means you can get almost the same performance for gesture recognition without needing access to the proprietary robot telemetry. You just strap some sensors to the surgeon's fingers and you're good to go.

Tom: And on the da Vinci suturing data, they couldn't compare against internal kinematics because they didn't have access. But they did compare against vision-only models. And the hand tracking data alone was competitive, and when they fused it with the video, they got the best results, an accuracy of zero point seven one.

Jane: So the external sensors aren't just a fallback. They're actually complementary to the video. The video can be occluded by blood or smoke or instruments, but the hand tracking doesn't care about that. It's always capturing the surgeon's motion.

Tom: And that's the key improvement MiDAS offers. It's not trying to replace the robot's internal sensors. It's providing an alternative that works on any platform, including ones where you have zero access to the internal data.

Jane: And they also showed that the RGB-D camera tracking, which uses a depth camera to watch the hands, works too, though not quite as well. It had more missed detections because of occlusions and the camera's limited field of view. But it's still a viable option for labs that don't want to put sensors on the surgeon.

Tom: So the takeaway here is that non-invasive sensing is a real path forward for surgical research. It's accurate, it's platform-agnostic, and it's affordable. Which brings us to our final segment, where we wrap this all up.

Conclusion: Tom: And that brings us to the end of our discussion on "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." Jane, I think we can both agree this is one of those papers that could really change how research gets done in this field.

Jane: Absolutely, Tom. The biggest takeaway for me is the accessibility. MiDAS gives labs a way to collect rich, synchronized, multimodal data from surgical robots without needing the manufacturer's cooperation. That's a huge barrier removed.

Tom: And they're not just talking about it. They released the code, the hardware designs, and two datasets, including that novel hernia repair suturing dataset on realistic tissue models. That's a concrete contribution that other researchers can build on immediately.

Jane: And the validation is solid. They showed the external sensors closely approximate the robot's internal kinematics, and they demonstrated that those external signals can power gesture recognition models just as well as the proprietary data. That's a strong proof of concept.

Tom: For the wider world, this could accelerate research into surgical training, skill assessment, and even patient safety. If we can automatically detect when a surgeon is struggling or when an error is about to happen, that could save lives.

Jane: And because MiDAS is platform-agnostic, it can be deployed on any robot, from the research-grade Raven-II to the clinical da Vinci Xi. That means the findings from this research could translate directly into the operating room.

Tom: Well said, Jane. We're going to say goodbye to "MiDAS" now, but we're definitely going to be watching for follow-up work from this team. Thanks for joining us, and we'll see you next time on the arXiv channel.

Keshara Weerasinghe, Seyed Hamid Reza Roodabeh, Andrew Hawkins, Zhaomeng Zhang, Zachary Schrader, Homa Alemzadeh

University of Virginia · University of Virginia Health System

cs.RO, cs.CV, cs.LG

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 29 pages, 17 figures

Project page: https://uva-dsa.github.io/MiDAS

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 89/100

Key concepts

MiDAS
Multimodal Data Acquisition System. This system captures data from multiple sources—surgeon fingers via electromagnetic trackers, depth cameras watching hands, force sensors on foot pedals, and video feed—and syncs them in real time. It works without touching the robot's internal software.
Platform-Agnostic
The MiDAS system is platform-agnostic because it can be used on different robots, such as the open-source Raven-II or the clinical da Vinci Xi. This means research findings are not tied to one specific proprietary robot and can be applied across various surgical platforms.
Gesture Recognition Models
These are state-of-the-art models, like MTRSAP and MS-TCN++, used to analyze the external hand tracking data from surgeons. The goal is to recognize specific surgical actions, such as 'approach peg' or 'make a C loop,' without relying on internal robot telemetry.
Hernia Repair Dataset
A new dataset collected during a surgical bootcamp featuring surgeons performing hernia repair suturing on high-fidelity tissue models made from real porcine tissue. This dataset includes annotated data with eight different suturing gestures.

Terminology

Summary

Summary

This paper introduces MiDAS, an open-source, platform-agnostic system for multimodal data acquisition in robot-assisted minimally invasive surgery (RMIS). The system enables time-synchronized, non-invasive data collection across surgical robotic platforms without requiring proprietary robot interfaces. MiDAS integrates electromagnetic hand tracking, RGB-D hand tracking, foot pedal sensing, and surgical video capture. The authors validated MiDAS on the open-source Raven-II and the clinical da Vinci Xi platforms by collecting multimodal datasets of peg transfer and hernia repair suturing tasks performed by surgical residents. Correlation analysis and downstream gesture recognition experiments were conducted, demonstrating that external hand and foot sensing closely approximated internal robot kinematics, and non-invasive motion signals achieved gesture recognition performance comparable to robot kinematics. The system and annotated datasets are publicly released, including the first multimodal dataset capturing hernia repair suturing on high-fidelity simulation models.

The paper states: We introduce MiDAS, an open-source, platform-agnostic system enabling time-synchronized, non-invasive multimodal data acquisition across surgical robotic platforms. The methods section notes: MiDAS integrates electromagnetic and RGB-D hand tracking, foot pedal sensing, and surgical video capturing without requiring proprietary robot interfaces. The results indicate: External hand and foot sensing closely approximated internal robot kinematics, and non-invasive motion signals achieved gesture recognition performance comparable to robot kinematics. The conclusions state: MiDAS enables reproducible multimodal RMIS data collection and is released with annotated datasets, including the first multimodal dataset capturing hernia repair suturing on high-fidelity simulation models.

The paper's contributions are fourfold: (1) presenting MiDAS as an open-source system that can be non-invasively integrated with standard teleoperated robotic surgery systems to collect fully synchronized data in real-time; (2) proposing new alternative modalities to proprietary robot kinematics, including surgeon hand and foot motion kinematics captured using electromagnetic hand tracking sensors, depth cameras, and force sensors, demonstrating they closely approximate surgeon and robot kinematics and can be effectively used for tasks such as activity recognition; (3) evaluating MiDAS by collecting data from two RMIS tasks—standard Peg Transfer on Raven-II and a novel multimodal dataset of hernia repair suturing on realistic KindHeart tissue models using the da Vinci Xi, performed during a training bootcamp at the University of Virginia Hospital—and identifying effective modality combinations for multimodal gesture recognition by benchmarking state-of-the-art Transformer and CNN models; (4) publicly releasing both the data collection system and the resulting multimodal, multi-platform datasets with gesture annotations.

The MiDAS architecture employs a client-server design: "A central server coordinates multiple client processes that stream heterogeneous data (e.g., kinematics, video) as shown in Figure 10 at different rates, implemented via Python multiprocessing workers to enable concurrent, real-time ingestion and buffering." The server provides a GUI for entering study metadata, selecting modalities, and monitoring system status. All streams are stamped on arrival at the MiDAS server with a Unix epoch timestamp, providing a unified, server-side time base. The estimated cost for a single surgeon-console MiDAS setup was approximately 8,000 USD, including the NDI trakSTAR electromagnetic tracking unit (4,850), Blackmagic SDI recorders (400), a ZED Mini RGB-D camera (400), a custom pedal sensing system (150), and a laptop workstation (2,000). The data recording usage rate is approximately 0.4 GB/minute when capturing at 30 Hz with 1080p stereoscopic video.

For electromagnetic hand tracking, the system employs the NDI trakSTAR device with four miniature 6-DoF sensors: Two sensors are mounted on each MTM control at the thumb and middle-finger contact pads (Figure 1), preserving the console hardware while measuring fine manipulator motions. The system logs 3D position and orientation for each sensor at 270 Hz. Raw sensor data are mapped from the electromagnetic tracker frame to robot reference frames via a calibrated rigid-body transformation augmented with a learned residual modeled using a Multi-Layer Perceptron (MLP). The MLP models have 2 hidden layers with 16 neurons, use ReLU activation, and are trained with root mean squared error objective, L2 weight decay, 50% dropout, Adam optimizer with learning rate 0.005, over 50 epochs.

For RGB-D hand tracking, a ZED Mini stereo camera is mounted above the surgeon console, providing synchronized 720p color video and dense depth maps at 30 Hz. Hand keypoints are extracted using Mediapipe's hand landmark detection library, and 3D positions are computed using the aligned RGB and depth frames. The keypoints are transformed into the MTM base frame using either a rigid body transformation from hand-eye calibration or an end-to-end MLP transformation. The paper notes: Limited camera field-of-view (FoV) and occlusions cause intermittent keypoint detection gaps. We fill gaps of at most one second using cubic spline interpolation and longer gaps using forward-filling.

For video capture, the system uses Open Broadcaster Software (OBS) as a unified capture layer. On da Vinci Xi, stereoscopic SDI output is captured using Blackmagic SDI recorders, while on Raven-II, stereo video is recorded using a ZED camera. OBS records synchronized, high-definition stereo streams at 30 Hz, controlled programmatically via the OBS WebSocket API.

The Pedal Sensing System (PSS) captures surgeon foot pedal interactions using an Arduino-class microcontroller and thin-film force-sensitive resistors (FSRs) affixed directly to console pedal surfaces. PSS supports up to nine pedals and records both binary pedal states and timestamped analog signals at 30 Hz. Each FSR is mounted in a voltage-divider with a known series resistor, and a brief per-pedal calibration establishes voltage thresholds for robust binary actuation detection, further refined through data-driven optimization.

Two datasets were collected. The Raven-II dataset consists of 15 Peg Transfer trials, resulting in 36 minutes of data with gestures annotated using the DESK taxonomy, involving 7 gesture classes (Approach peg, Align & grasp, Lift peg, Transfer peg – Get together, Transfer peg – Exchange, Approach pole, Align & place) with 345 total samples. Ground-truth kinematic data including MTM and PSM position, orientation, velocity, grasper angle, and clutch pedal presses were available from the open-source platform. The da Vinci Xi dataset was collected during a robotic surgery training bootcamp at the University of Virginia hospital, with 40 participants (mostly 2nd and 3rd year surgical residents) performing 7 trials of inguinal hernia repair and 10 trials of ventral hernia repair, totaling 3.5 hours of annotated suturing segments. The suturing taxonomy includes 8 gestures: Orient Needle, Target Needle, Push Needle Through Tissue, Pull Needle out of Tissue, Reach Suture, Pull Suture, Make C Loop, and Square Knot and Cinch, with 1,724 total samples across 212 minutes. The taxonomy was developed in collaboration with an expert robotic surgeon and iteratively refined and validated.

Survey responses from bootcamp participants indicated strong acceptance: "92.3% of participants reporting overall satisfaction with the training experience and 86.1% rating the KindHeart tissue models as realistic. Importantly, 80% of participants disagreed or strongly disagreed that the sensors negatively affected their psychomotor performance."

For validation, the authors performed cross-modal correlation analysis comparing EmHT and HandKP data against ground-truth internal MTM and PSM kinematics on Raven-II. For positional trajectories, EmHT-MTM and EmHT-PSM mean Cosine Similarity (CoS) exceeded 0.8 across all axes, with Normalized Root Mean Square Error (NRMSE) below 24.4%. HandKP achieved mean CoS exceeding 0.7 for all axes with superior 0.88-0.89 Z-axis alignment. For orientation, EmHT-MTM achieved CoS of 0.90 for Yaw, though Pitch and Roll showed moderate alignment (0.61 and 0.65). EmHT-PSM orientation alignment was lower, with CoS values ranging from 0.32 (Roll) to 0.44 (Pitch), attributed to the complex control loop and Raven-II's controller errors. For grasper angle, EmHT sensors achieved moderate agreement with both MTM and PSM grasper angles (IoU ≈ 0.53-0.54, Accuracy ≈ 0.80-0.81).

For pedal sensing validation, the PSS was compared against ground-truth pedal data from both platforms. On Raven-II, PSS achieved F1 of 0.85, precision of 0.78, recall of 0.95, IoU of 0.74, and lag of 166.67 ms. On da Vinci Xi, PSS achieved F1 of 0.78, precision of 0.69, recall of 0.86, IoU of 0.67, and lag of 133.33 ms. The paper notes: "The lower precision (0.69) arises primarily from the intrinsic high-sensitivity design of the pedal sensing system, which is biased toward detecting all true presses (i.e., favoring recall) and therefore produces occasional false positives."

For downstream gesture recognition, the authors benchmarked two state-of-the-art temporal models, MTRSAP and MS-TCN++, on both datasets. On Raven-II Peg Transfer, MTM/PSM kinematics achieved F1-scores of 0.87-0.88. Notably, using non-invasive EmHT achieves comparable performance to MTM/PSM, with MTRSAP model gaining an F1-score of 0.86 and accuracy of 0.87, closely matching baseline MTM/PSM models. HandKP performed suboptimally (F1-score of 0.4) due to suboptimal pose approximation and missed hand detections. On da Vinci Xi Suturing, EmHT performed comparably to vision-only models and improved accuracy when fused with them: MTRSAP EmHT+Image (DINOv2) achieving Acc of 0.71 and F1-score of 0.70, outperforming each single modality. Image-only models underperformed due to limited data, with DINOv2 features outperforming ResNet-50.

The paper also reports ablation studies. For feature ablation on Raven-II, incorporating clutch state information led to notable performance gains for EmHT (F1 improved from 0.77 to 0.86). Removing grasper state caused macro-F1 to drop significantly from 0.78 to 0.47 in MTM performance and 0.75 to 0.58 in PSM performance. Incorporating velocity consistently yielded substantial gains across kinematic modalities. For visual feature adaptation, pretraining ResNet-50 on the in-domain DESK dataset yielded 15-30% relative improvement, and replacing ResNet-50 with DINOv2 features provided further gains.

The paper concludes: "By releasing the MiDAS system, multimodal datasets, and baseline benchmarks, we aim to support reproducible, cross-platform research, reduce reliance on closed robotic systems, and accelerate progress in multimodal learning for surgical activity understanding." The work was partially supported by research awards from the National Science Foundation (CNS-2146295) and the Commonwealth Cyber Initiative. The surgical bootcamp and associated data collection were approved by the University of Virginia Institutional Review Board under protocol IRB-SBS 6724, and all participants provided informed consent.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:

Improvement: I will implement a cross-modal attention fusion mechanism that combines electromagnetic hand tracking (EmHT) signals with vision features (DINOv2), as demonstrated by the MTRSAP model achieving 0.71 accuracy and 0.70 F1-score on da Vinci Xi suturing—outperforming single-modality baselines by 2-4% absolute.

Capability: The improved system can recognize 8 distinct suturing gestures (needle orientation, targeting, pushing through tissue, pulling, suture reaching, pulling suture, C-loop formation, square knot cinching) in real-time from non-invasive sensors, even when video quality degrades due to occlusion or lens contamination.

Abstract

Background: Robot-assisted minimally invasive surgery (RMIS) research increasingly relies on multimodal data, yet access to proprietary robot telemetry remains a major barrier. We introduce MiDAS, an open-source, platform-agnostic system enabling time-synchronized, non-invasive multimodal data acquisition across surgical robotic platforms. Methods: MiDAS integrates electromagnetic and RGB-D hand tracking, foot pedal sensing, and surgical video capturing without requiring proprietary robot interfaces. We validated MiDAS on the open-source Raven-II and the clinical da Vinci Xi by collecting multimodal datasets of peg transfer and hernia repair suturing tasks performed by surgical residents. Correlation analysis and downstream gesture recognition experiments were conducted. Results: External hand and foot sensing closely approximated internal robot kinematics and non-invasive motion signals achieved gesture recognition performance comparable to proprietary telemetry. Conclusion: MiDAS enables reproducible multimodal RMIS data collection and is released with annotated datasets, including the first multimodal dataset capturing hernia repair suturing on high-fidelity simulation models.

Sources

Related papers