Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition
cs.CV
Submitted: 2026-07-22
Updated: 2026-10-06
Project page: https://marginlab.github.io/AI-for-healthcare/lavifit
Terminology
Sources
- CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
- SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
- fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
- EmbeddingGemma: Powerful and Lightweight Text Representations
- DINOv2: Learning Robust Visual Features without Supervision
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
- CholecSeg8k: A Semantic Segmentation Dataset for Laparoscopic Cholecystectomy Based on Cholec80
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- SurgMAE: Masked Autoencoders for Long Surgical Video Analysis
- Whether and When does Endoscopy Domain Pretraining Make Sense?
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
- Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets
- Surgical Triplet Recognition via Diffusion Model
- TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models
- Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
- Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
- Latent Action Learning Requires Supervision in the Presence of Distractors
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models