A Surgical Foundation Model Reveals Task-Dependent Label Efficiency
cs.CV, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/CAMMA-public/SURGE
Terminology
Sources
- The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark
- 2017 Robotic Instrument Segmentation Challenge
- SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection
- Surgical Anatomy Recognition with Context Learning using Foundation Representations
- ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection
- Mugs: A Multi-Granular Self-Supervised Learning Framework
- iBOT: Image BERT Pre-Training with Online Tokenizer
- DINOv3
- CliPPER: Contextual Video-Language Pretraining on Long-form Intraoperative Surgical Procedures for Event Recognition
- General surgery vision transformer: A video pre-trained foundation model for general surgery
- Robust Medical Instrument Segmentation Challenge 2019
- SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
- Surgical Visual Understanding (SurgVU) Dataset
- The SARAS Endoscopic Surgeon Action Detection (ESAD) dataset: Challenges and methods
- The SAGES Critical View of Safety Challenge: A Global Benchmark for AI-Assisted Surgical Quality Assessment
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Spreading vectors for similarity search
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models