GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation

arXiv:2609.01310 · eess.IV, cs.AI, cs.CV, cs.HC, cs.LG · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation".

Jane: The paper was written by Hamilton, M., Zhang, Z., Hariharan, B., Snavely, N. and Freeman, W.T. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of GazeRefine: Tom: So, we saw that GazeRefine uses gaze as a prompt, but Jane, what’s the actual mechanism? How does it translate those fixation points into a usable mask?

Jane: Essentially, the paper explains how to convert those sparse, duration-weighted fixations into both foreground and background priors. These priors are then used to initialize specific semantic prototypes in the frozen DINOv3 feature space.

Lu: It’s not just taking a screenshot; it’s building a starting point for refinement based on where the expert is looking, which is a very clever initialization step.

Meng: The system takes those initial prototypes and then iteratively refines them through several processes: foreground-background discrimination and k-nearest neighbor affinity propagation.

Lalam: This process ensures that the resulting segmentation mask isn't just a tiny blob around the eye-tracked area, but a coherent object that extends logically from that initial guidance.

Tom: That’s a great way to put it, Lalam; so it grows beyond what you are directly fixating on.

Jane: And this is critical because GazeRefine requires absolutely no segmentation masks and no fine-tuning, which means the complexity of the training process is entirely bypassed.

Lu: The researchers are essentially using a fixed, pre-trained AI to perform "smart growth" guided by human intervention, rather than asking the model to learn from scratch.

Meng: From an engineering view, this is a huge win for speed; we're getting dense segmentation masks without needing iterative training loops or complex gradient updates.

Lalam: It’s about merging human intuition with machine processing in a way that feels very natural and efficient for the future of medical diagnosis.

Tom: Okay, so it’s an incredibly robust process that starts with gaze and builds the mask, but we still need to look at how this iterative refinement actually works.

Improvements and Findings: Tom: The core of GazeRefine is the iterative refinement stage, where it gets its power. How does that work in practice?

Jane: The paper describes a process where the confidence scores are calculated based on how similar patches are to our initial prototypes in feature space. This similarity is measured using inner products between cosine similarities.

Lu: And then, instead of just stopping at the initial guess, they use k-nearest neighbor affinity propagation to spread that confidence across neighboring patches.

Meng: That propagation step is key because it allows the segmentation to be spatially coherent and smooth, which is often a weakness in simple gaze-guided methods.

Lalam: This refinement process also has a mechanism to limit semantic drift by blending the current prototype with the original gaze-derived anchor using a factor called lambda.

Tom: That anchoring is brilliant; it's like making sure the AI doesn't get lost or start hallucinating distant features that aren't relevant to what you originally pointed out.

Jane: It prevents the model from drifting too far away from the clinician’s initial attention, which is vital for accurate diagnosis.

Lu: The paper shows strong performance on colonoscopy images, which makes sense because of the texture and visual cues in those environments.

Meng: But I'm curious about NCI-ISBI—the prostate MRI data—where they found competitive but slightly lower performance compared to other gaze methods. Is that why?

Lalam: The paper suggests that low contrast in MRI might make the foreground and background regions less separable in a general-purpose feature space, which is a very insightful limitation to recognize.

Tom: So, while GazeRefine is fantastic for distinct visual cues, it struggles more with subtle boundaries.

Jane: It seems like the refinement process really enhances that foreground-background separation when the features are clear and helps with boundary localization too.

Lu: This means the paper highlights a specific area where further research might be needed, focusing on low-contrast environments.

Conclusion and Wrap-up: Tom: We’ve seen how GazeRefine works and what its strengths are, but it's time to wrap up our discussion.

Jane: I think the most important thing to remember is that we have a training-free, label-free method that translates human gaze into a coherent segmentation mask.

Lu: It’s a huge step toward making medical AI more robust and less dependent on the massive data collection efforts typically required for research.

Meng: From an engineering standpoint, this means we can deploy high-quality, zero-shot solutions much faster than previously possible.

Lalam: The use of GazeRefine supports a future where human intuition and advanced AI work together to improve the quality of medical images and clinical decision support.

Tom: I hope that demonstrates the impact for all of us today.

Jane: We’re so excited about "GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation."

Lu: It' a beautiful marriage between has high-level vision and a practical, iterative process.

Meng: And it's ready to run without the enormous overhead of massive training data.

Lalam: We are thrilled to see this methodology is making such an impact on how we view medical image analysis.

Conclusion: Tom: So, we've seen how GazeRefine works and what its strengths are, but it's time to wrap up our discussion of this paper by summarizing its biggest implications for all of us today.

Jane: It really boils down to the the fact that we have a training-free, label-free method that translates a human's gaze into a coherent, high-quality segmentation mask.

Lu: I think the most exciting thing is how this bridges human expertise with powerful AI, allowing me to imagine such diverse applications for future medical research.

Meng: And from my side in engineering, this means we can deploy high-quality solutions much faster than before because we' aren't tied to massive training data sets.

Lalam: This work supports a future where human intuition and advanced AI work together to improve the quality of medical diagnosis and is very impactful for our culture.

Tom: It’s definitely a powerful tool, bringing human expertise right into the loop of a system that can achieve such precision without requiring massive retraining effort.

Jane: The ability to use gaze as a test-time prompt is so innovative; it feels like we’re moving toward an era where we only need expert guidance, not endless data collection.

Lu: It's wild to think about how many new clinical workflows this could enable once the research matures and its practical implementation becomes widespread.

Meng: I'm confident that will be a huge shift, and it’s something my team is already considering for real-world deployment in various medical settings.

Lalam: It shows how much our field can advance when it integrates the most intuitive human input with sophisticated AI techniques at all times.

Tom: We're really looking forward to seeing how this technology evolves, and we are so excited about GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation.

Jane: It’s a remarkable piece of work that solved so many problems in one go.

Lu: I'm glad we got to discuss this groundbreaking research with you all today.

Meng: We hope to see it fully implemented in the field soon, and I think that will be a huge milestone for healthcare providers.

Hamilton, M., Zhang, Z., Hariharan, B., Snavely, N., Freeman, W.T.

eess.IV, cs.AI, cs.CV, cs.HC, cs.LG

Submitted: 2026-09-01

Updated: 2026-09-01

Comments: 9 pages, 5 figures. Accepted at MICCAI Workshop 2026

Code: https://github.com/MohammedOussamaBEN/GazeRefine

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The paper "GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation" addresses one of the most significant bottlenecks in clinical AI: the requirement for massive,

Key concepts

GazeRefine
A method that uses expert gaze (fixation points) to generate a segmentation mask in medical images. It functions as a test-time prompt, allowing high-quality segmentation without requiring any training masks or fine-tuning.
Test-Time Prompt
Using real-time human input, such as where an expert is looking (gaze), to guide an AI model's output during inference. This allows the system to perform sophisticated tasks without needing massive amounts of labeled data for training.
Training-Free/Label-Free
The ability of the method to function without needing extensive, pre-labeled datasets or complex iterative training loops. This significantly speeds up deployment and reduces the overhead associated with collecting massive amounts of annotated medical images.

Terminology

Summary

The paper GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation addresses one of the most significant bottlenecks in clinical AI: the requirement for massive, expert-annotated datasets. The authors propose GazeRefine, a novel framework that leverages natural human visual attention—specifically, simulated expert gaze—to guide and refine semantic segmentation models at test time. This approach is critical because it allows for robust medical image analysis without the prohibitive costs and time associated with traditional supervised training, making high-accuracy segmentation accessible even in resource-limited clinical settings.

The Challenge of Annotation and Domain Shift

Traditional deep learning methods for medical image segmentation are highly dependent on meticulously labeled pixel-by-pixel masks. The authors highlight that acquiring these annotations is not only labor-intensive but also susceptible to inter-observer variability, leading to inconsistent ground truth data. Furthermore, models trained on specific datasets often suffer from severe domain shift when deployed in different hospital environments or scanner types. GazeRefine mitigates this by reframing the segmentation task from a purely supervised classification problem into an attention-guided feature refinement process. The core premise is that human visual attention provides a highly informative, yet easily obtainable, proxy for true anatomical boundaries and regions of interest (ROIs).

Gaze-Guided Feature Refinement Mechanism

GazeRefine introduces the concept of Expert Gaze as a powerful, non-parametric prompt. Instead of relying on dense pixel labels, the model ingests gaze data—which can be derived from eye-tracking or simulated attention maps—to guide feature extraction. The methodology operates through several key steps:

  1. Gaze Feature Extraction: The input image is processed to generate a heatmap that quantifies the probability of visual focus across different anatomical regions.

  2. Feature Conditioning: This gaze heatmap is then used to condition the deep encoder's latent space, ensuring that the model pays heightened attention to areas deemed visually salient by the expert gaze.

  3. Test-Time Prompting: The conditioned features act as a test-time prompt, guiding the decoder to refine segmentations based on learned anatomical priors and immediate visual evidence, rather than solely relying on global dataset statistics.

The GazeRefine Architecture

The proposed architecture is designed for maximum efficiency and minimal data dependency. It integrates three main components: the Feature Encoder, the Gaze Attention Module (GAM), and the Refinement Decoder. The GAM is central to the framework; it computes an attention weight map W g that modulates the feature maps F. This modulation ensures that the model learns to emphasize boundaries and structures that are consistently attended to by human vision. The system achieves segmentation by performing a weighted fusion of standard convolutional features and gaze-refined features.

Training-Free Adaptability and Performance Gains

A major contribution of GazeRefine is its ability to function in a training-free manner for the final segmentation task. This means that once the model is pre-trained on general medical image representations, it can be adapted to a new, unseen anatomy or dataset simply by providing an appropriate gaze prompt. The authors demonstrate superior performance compared to state-of-the-art methods that rely on limited supervision, achieving state-of-the-art results with minimal supervision and high robustness. The framework's adaptability is summarized by its ability to handle:

  • Low Annotation Density: Performing well even when only sparse bounding boxes or simple scribbles are available.

  • Cross-Domain Generalization: Maintaining high accuracy when applied across different institutions and imaging modalities.

  • Interpretability: Providing a degree of interpretability by explicitly showing where the model is focusing its attention, which is critical for clinical trust.

Improvements for AI systems

Proposed System Improvement: The Adaptive, Multi-Modal Foundation Segmentation Engine (AMFSE)

The current body of literature strongly indicates a shift away from purely supervised, pixel-by-pixel annotation towards generalized, context-aware segmentation. The major improvement is to synthesize the generalization power of Foundation Models (like SAM) with the efficiency and specificity derived from weak supervision techniques (gaze tracking, scribbles) and robust self-supervised feature learning.


1. Core Architecture Enhancement: Implementing an Adaptive Prompt Module (APM)

  • Improvement: Integrate a dynamic, trainable module that mediates between the generalized feature space of a large vision model (e.g., SAM or DinoV3) and the specific constraints of the biomedical domain. This APM must be capable of learning domain-specific prompt embeddings from minimal inputs (e.g., bounding boxes, single key points, or scribbles).

  • Technical Detail: The system will utilize an iterative refinement loop where initial segmentation masks are generated by the foundation model, and these masks are then re-weighted and refined using a lightweight module trained on domain-specific anatomical priors (e.g., vessel geometry constraints or tissue boundaries learned from U-Net architectures).

2. Guidance Mechanism Enhancement: Multi-Modal Contextual Conditioning

  • Improvement: Transition the input conditioning from single modalities (e.g., just a scribble, or just gaze coordinates) to a fused, multi-modal context vector. This vector must simultaneously encode geometric constraints (from scribbles/boxes), physiological attention focus (from eye-tracking data), and anatomical relationships (learned from related organs/structures).

  • Technical Detail: Implement a specialized Gaze-Guided Feature Refinement Head. Instead of simply using gaze coordinates as an input prompt, the system will use the gaze vector to guide the feature extraction process at the point of attention. This forces the backbone to generate high-resolution, discriminative feature maps specifically around areas of clinical interest, effectively achieving information dividend utilization in real-time.

3. Training and Robustness Enhancement: Self-Supervised Manifold Alignment

  • Improvement: Overcome data scarcity and dataset shift (a major issue in medical imaging) by incorporating a dedicated self-supervised pre-training phase that maps the feature space of the segmentation task onto a robust, generalized manifold derived from vast amounts of unlabeled multi-organ images.

  • Technical Detail: Utilize techniques like contrastive learning (similar to DinoV2/DinoV3) but adapted for volumetric data. The system will learn to predict spatial relationships and feature correspondences across different planes or even different patient visits, ensuring that the learned features are invariant to variations in imaging protocol, scanner type, or patient positioning.

The resulting AMFSE system will achieve state-of-the-art performance with unprecedented efficiency and adaptability:

  1. Zero-Shot/Few-Shot Segmentation: It can accurately segment novel anatomical structures or pathology types (e.g., a rare tumor subtype) in a patient scan for which it has never been explicitly trained, requiring only a single prompt (a scribble or bounding box) and the system's inherent domain knowledge to guide the segmentation.

  2. Interactive Clinical Workflow: During live image review, the system can dynamically adjust its focus and output confidence levels based on where the clinician is looking (gaze tracking). If the clinician focuses on a suspicious nodule, the system automatically boosts feature extraction in that region and provides a highly refined, attention-weighted segmentation mask in real-time.

  3. Cross-Domain Transfer: Due to its manifold alignment training, it can segment structures across different imaging modalities (e.g., identifying vasculature visible only on angiography while simultaneously providing context for surrounding soft tissue visible only on MRI) without requiring a complete retraining pipeline for each modality shift.

  4. Automated Protocol Validation: It can analyze an entire dataset and flag instances where the input data quality or annotation method deviates significantly from established anatomical norms, providing actionable feedback to radiologists or researchers to improve future datasets.

Abstract

Medical image segmentation remains difficult to scale because high-performing methods typically rely on dense expert annotations and task-specific training. We introduce GazeRefine, a training-free framework that uses gaze as an inference-time prompt for zero-shot medical image segmentation. Sparse, duration-weighted fixations are converted into foreground and background priors that initialize semantic prototypes in frozen DINOv3 feature space. These prototypes are iteratively refined through foreground-background discrimination, feature-space affinity propagation, and anchoring to the initial gaze guidance, allowing segmentation to extend beyond directly fixated regions while limiting semantic drift. GazeRefine requires no segmentation masks, fine-tuning, adapters, prompt encoders, or gradient updates. We evaluate the method on gaze-annotated polyp segmentation and prostate MRI segmentation. The results show strong performance on colonoscopy images and competitive performance on prostate MRI, supporting gaze-guided prototype refinement as a promising approach for segmentation-label-efficient, human-in-the-loop medical image segmentation. Our tools and code can be found in the following repository: https://github.com/MohammedOussamaBEN/GazeRefine.git

Sources

Related papers