MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery
summary
In short
The episode discusses 'MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery,' a system that captures synchronized data from surgeons, robots, and video without needing proprietary robot access. The team validated the system on two different robots and released the code, hardware designs, and datasets to allow researchers to study surgical skill more broadly.
Key concepts
- MiDAS
- Multimodal Data Acquisition System. This system captures data from multiple sources—surgeon fingers via electromagnetic trackers, depth cameras watching hands, force sensors on foot pedals, and video feed—and syncs them in real time. It works without touching the robot's internal software.
- Platform-Agnostic
- The MiDAS system is platform-agnostic because it can be used on different robots, such as the open-source Raven-II or the clinical da Vinci Xi. This means research findings are not tied to one specific proprietary robot and can be applied across various surgical platforms.
- Gesture Recognition Models
- These are state-of-the-art models, like MTRSAP and MS-TCN++, used to analyze the external hand tracking data from surgeons. The goal is to recognize specific surgical actions, such as 'approach peg' or 'make a C loop,' without relying on internal robot telemetry.
- Hernia Repair Dataset
- A new dataset collected during a surgical bootcamp featuring surgeons performing hernia repair suturing on high-fidelity tissue models made from real porcine tissue. This dataset includes annotated data with eight different suturing gestures.
Terminology used across episodes
This episode discusses
- MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery · Paper Radio
- SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
- MS-TCRNet: Multi-Stage Temporal Convolutional Recurrent Networks for Action Segmentation Using Sensor-Augmented Kinematics
- DINOv2: Learning Robust Visual Features without Supervision
- ASFormer: Transformer for Action Segmentation
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- MediaPipe: A Framework for Building Perception Pipelines
The paper
MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery · Read on arXiv
Keshara Weerasinghe, Seyed Hamid Reza Roodabeh, Andrew Hawkins, Zhaomeng Zhang, Zachary Schrader, Homa Alemzadeh
University of Virginia · University of Virginia Health System
Background: Robot-assisted minimally invasive surgery (RMIS) research increasingly relies on multimodal data, yet access to proprietary robot telemetry remains a major barrier. We introduce MiDAS, an open-source, platform-agnostic system enabling time-synchronized, non-invasive multimodal data acquisition across surgical robotic platforms. Methods: MiDAS integrates electromagnetic and RGB-D hand tracking, foot pedal sensing, and surgical video capturing without requiring proprietary robot interfaces. We validated MiDAS on the open-source Raven-II and the clinical da Vinci Xi by collecting multimodal datasets of peg transfer and hernia repair suturing tasks performed by surgical residents. Correlation analysis and downstream gesture recognition experiments were conducted. Results: External hand and foot sensing closely approximated internal robot kinematics and non-invasive motion signals achieved gesture recognition performance comparable to proprietary telemetry. Conclusion: MiDAS enables reproducible multimodal RMIS data collection and is released with annotated datasets, including the first multimodal dataset capturing hernia repair suturing on high-fidelity simulation models.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery".
Jane: The paper was written by Keshara Weerasinghe, Seyed Hamid Reza Roodabeh, Andrew Hawkins, Zhaomeng Zhang, Zachary Schrader et al. from University of Virginia and University of Virginia Health System.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the arXiv channel, everyone. I'm Tom, and as always, Jane is here with me. Today we're looking at a paper that's going to make a lot of robotics labs very happy. It's called "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery."
Jane: And Tom, I have to say, the first thing that struck me is the team behind this. It's from the University of Virginia, and they've got engineers working side by side with a thoracic and cardiovascular surgeon. That mix is exactly what you need when you're trying to solve a problem in the operating room.
Tom: Absolutely. And the problem they're tackling is one that anyone who's tried to do research on surgical robots knows all too well. The robots themselves, like the da Vinci system, are these closed, proprietary boxes. You can't just plug in and get the data you need.
Jane: Right. If you want to study how a surgeon moves, or how they press the pedals, or what the robot's arms are doing, you're usually stuck. The manufacturer doesn't give you easy access to that internal telemetry. So this team built their own system to capture it from the outside, without touching the robot's software at all.
Tom: And that's the "MiDAS" part. It stands for Multimodal Data Acquisition System. They're using electromagnetic trackers on the surgeon's fingers, a depth camera watching their hands, force sensors on the foot pedals, and the surgical video feed. All of it synced together in real time.
Jane: What I love is that they validated it on two very different robots. They tested it on the Raven-II, which is this open-source research robot, and then they took it to a clinical da Vinci Xi system during a real surgical training bootcamp. That's a huge deal because it proves the system isn't tied to one platform.
Tom: And they're giving it all away. The code, the hardware designs, and the datasets they collected, including a brand new one of surgeons doing hernia repair on realistic tissue models. That's a gift to the research community.
Jane: It really is. And it means that labs that don't have access to expensive, proprietary robots can still do meaningful research on surgical skill and safety. They just need a MiDAS setup and a robot to practice on.
Tom: So stick around, because in the next segment we're going to dig into what they actually collected and why that hernia repair dataset is such a big deal.
Summary: Tom: Welcome back. We're still talking about "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." And Jane, I want to get into the actual data they collected, because there's a lot here.
Jane: There really is. They ran two big experiments. On the Raven-II, they had people do the classic peg transfer task, moving little rings from one peg to another. That gave them fifteen trials and about thirty-six minutes of data, all annotated with seven different gestures like "approach peg" and "transfer peg."
Tom: And that's the dry-lab stuff. But then they did something much more ambitious. They went to a surgical bootcamp at the University of Virginia hospital and recorded surgeons doing actual hernia repair suturing on these high-fidelity tissue models made by a company called KindHeart.
Jane: Those models are made from real porcine tissue, so they feel much closer to real surgery than a plastic trainer. They got seventeen trials, over two hundred minutes of annotated data, with eight different suturing gestures. That's a lot of careful labeling work.
Tom: And the gesture taxonomy itself is interesting. They didn't just make it up. They built it with an expert surgeon, and it captures things like "orient needle," "push needle through tissue," and "make a C loop." These are the actual steps a surgeon goes through when closing a hernia.
Jane: What I find really impressive is the validation they did. They compared their external sensors against the robot's own internal kinematics on the Raven-II. The electromagnetic hand trackers matched the robot's motion really well, with cosine similarity above zero point eight on the position data.
Tom: And the foot pedal sensors, they compared those against the ground truth too. They got an F1 score of zero point eight five on the Raven-II and zero point seven eight on the da Vinci. So the system is actually capturing what's happening, not just guessing.
Jane: Exactly. And that's crucial because it means researchers can trust this data. They can use it to train models for gesture recognition, for skill assessment, even for detecting errors during surgery. And they don't need to hack into the robot to do it.
Tom: Which brings us to the next question. How well do these external signals actually work for recognizing what the surgeon is doing? That's what we're going to dig into next.
Improvements and Experiments: Tom: We're back with "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." And Jane, we just talked about the data collection. Now I want to talk about what they did with it, because they ran some serious experiments.
Jane: They did. They took their external hand tracking data and fed it into two state-of-the-art gesture recognition models. One is called MTRSAP, which is a transformer-based model, and the other is MS-TCN++, which is a temporal convolutional network.
Tom: And the results are pretty striking. On the peg transfer task, the electromagnetic hand tracking data performed almost as well as the robot's own internal kinematics. We're talking about an F1 score of zero point eight six versus zero point eight seven for the robot data. That's a tiny gap.
Jane: That's a huge finding. It means you can get almost the same performance for gesture recognition without needing access to the proprietary robot telemetry. You just strap some sensors to the surgeon's fingers and you're good to go.
Tom: And on the da Vinci suturing data, they couldn't compare against internal kinematics because they didn't have access. But they did compare against vision-only models. And the hand tracking data alone was competitive, and when they fused it with the video, they got the best results, an accuracy of zero point seven one.
Jane: So the external sensors aren't just a fallback. They're actually complementary to the video. The video can be occluded by blood or smoke or instruments, but the hand tracking doesn't care about that. It's always capturing the surgeon's motion.
Tom: And that's the key improvement MiDAS offers. It's not trying to replace the robot's internal sensors. It's providing an alternative that works on any platform, including ones where you have zero access to the internal data.
Jane: And they also showed that the RGB-D camera tracking, which uses a depth camera to watch the hands, works too, though not quite as well. It had more missed detections because of occlusions and the camera's limited field of view. But it's still a viable option for labs that don't want to put sensors on the surgeon.
Tom: So the takeaway here is that non-invasive sensing is a real path forward for surgical research. It's accurate, it's platform-agnostic, and it's affordable. Which brings us to our final segment, where we wrap this all up.
Conclusion: Tom: And that brings us to the end of our discussion on "MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery." Jane, I think we can both agree this is one of those papers that could really change how research gets done in this field.
Jane: Absolutely, Tom. The biggest takeaway for me is the accessibility. MiDAS gives labs a way to collect rich, synchronized, multimodal data from surgical robots without needing the manufacturer's cooperation. That's a huge barrier removed.
Tom: And they're not just talking about it. They released the code, the hardware designs, and two datasets, including that novel hernia repair suturing dataset on realistic tissue models. That's a concrete contribution that other researchers can build on immediately.
Jane: And the validation is solid. They showed the external sensors closely approximate the robot's internal kinematics, and they demonstrated that those external signals can power gesture recognition models just as well as the proprietary data. That's a strong proof of concept.
Tom: For the wider world, this could accelerate research into surgical training, skill assessment, and even patient safety. If we can automatically detect when a surgeon is struggling or when an error is about to happen, that could save lives.
Jane: And because MiDAS is platform-agnostic, it can be deployed on any robot, from the research-grade Raven-II to the clinical da Vinci Xi. That means the findings from this research could translate directly into the operating room.
Tom: Well said, Jane. We're going to say goodbye to "MiDAS" now, but we're definitely going to be watching for follow-up work from this team. Thanks for joining us, and we'll see you next time on the arXiv channel.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language