Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

arXiv:2603.06697 · cs.CV, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs".

Jane: The paper was written by Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu et al. from University of Georgia and University of Texas, Arlington and Massachusetts General Hospital, Harvard Medical School and New Jersey Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper with a title that just grabs you: "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs." Jane, I gotta say, that title is doing a lot of heavy lifting.

Jane: It really is, Tom. And it's from a team spanning the University of Georgia, UT Arlington, Harvard, and NJIT. The core idea is that when a radiologist reads a chest X-ray, they don't just stare at the whole image at once. Their eyes move in a specific sequence, hopping from one suspicious spot to another. That's the "gaze" part.

Tom: Right, and that sequence isn't random. It's the doctor's actual reasoning process made visible. The paper argues that most vision-language models, or VLMs, they take the image in, but then they "think" in text. They describe what they see and then reason about that description.

Jane: Exactly. And that's a bottleneck. Some visual information just doesn't translate well into words. So this team's idea is to use the radiologist's eye-tracking data as a kind of "thinking" supervision. They're teaching the model to follow the same visual path the expert did, step by step.

Tom: So instead of just saying "there's a shadow here," the model learns the *process* of finding that shadow. That's a pretty profound shift in how we think about training these models, isn't it?

Jane: It is. It's moving from teaching the model *what* to see, to teaching it *how* to look. And that's a much richer signal. The authors even call it "visual reasoning supervision." It's not just an attention map; it's a time-ordered record of evidence gathering.

Tom: And that temporal order is the key, right? We'll get into the details, but I'm already excited about the potential here. It's like giving the AI a mentor who shows it their work, not just the final answer.

Jane: And that mentor is a board-certified radiologist. The implications for training and for building trust in these systems are huge. Stick around, because we're going to dig into how they actually made this work.

Summary: Tom: So, Jane, we've got this great title. Now, what's the actual meat of "Thinking with Gaze"? Give me the quick version.

Jane: Okay, so they took a powerful VLM, specifically Qwen2 point 5-VL-7B, and they fine-tuned it on a dataset called MIMIC-EYE. That dataset has chest X-rays, the radiologist's report, and crucially, the synchronized eye-tracking data from when the radiologist was reading the image.

Tom: And that's the secret sauce. But how do you get a language model to "look" at something? It just predicts text tokens.

Jane: Right, so they did something clever. They added four special placeholder tokens to the model's vocabulary. Let's call them "gaze tokens." The model is trained to output these tokens at the start of its answer. Then, they take the internal hidden states at those token positions and use a small projection head to predict which patch of the image the radiologist was looking at.

Tom: So the model is literally being forced to predict the doctor's eye movements, in order, before it gives its final diagnosis. That's wild.

Jane: It is. And they do this in two stages. Stage one is all about learning this gaze prediction. Stage two is the actual classification task, where the model has to output a "yes" or "no" for fourteen different findings, like cardiomegaly or pneumonia. They also use LoRA, which is a clever trick to fine-tune the model without having to update all of its billions of parameters.

Tom: And the results? They compared their method, which they call "Original-Gaze," against a bunch of baselines. What happened?

Jane: On the in-domain MIMIC-EYE test set, their method hit an AUROC of ninety point one seven. That's a big jump from the supervised fine-tuning baseline, which was at eighty-seven point six zero. And it beat out other gaze-based methods, like just using a static heatmap.

Tom: So preserving the order of the gaze is important. It's not just about *where* they looked, but *when*.

Jane: Precisely. They even ran an ablation where they shuffled the gaze order, and performance dropped. That tells you the temporal sequence is a crucial part of the signal. It's not just a spatial prior; it's a temporal reasoning chain.

Tom: That's a really strong finding. It validates the whole premise of the paper. I'm curious to see how this holds up on data the model has never seen before. That's what we're going to talk about next.

Improvements: Tom: We're back with "Thinking with Gaze," and Jane just laid out the in-domain results. But the real test for any model like this is, does it work on new data? Does it generalize?

Jane: And that's exactly where this paper shines, Tom. They took their trained model and tested it zero-shot on three completely different chest X-ray datasets: CheXpert, RSNA, and SIIM-ACR. No fine-tuning on those datasets at all.

Tom: Zero-shot, meaning it has to figure it out from scratch. That's a tough test. And how did the gaze-trained model do?

Jane: It was the best performer on every single one of those benchmarks. For example, on the RSNA dataset, their "Original-Gaze" method got an F1 score of fifty-three point seven three, compared to forty-three point seven seven for the standard supervised fine-tuning. That's a massive improvement in a metric that's often hard to move.

Tom: Wow. So it's not just memorizing the MIMIC-EYE data. It's learning a transferable skill. It's learning *how* to look for evidence.

Jane: Exactly. And that's the key implication. The authors argue that by mimicking the radiologist's visual search pattern, the model is less likely to rely on dataset-specific shortcuts. It's learning a more robust, human-like strategy for gathering evidence.

Tom: So it's not just about being more accurate; it's about being more reliable. That's huge for clinical deployment, where you're going to see images from all sorts of different machines and settings.

Jane: Right. And it also has an interpretability angle. The model is producing a gaze sequence as part of its output. That means a clinician can look at that sequence and see, "Oh, the model looked at the left lung base first, then the cardiac silhouette." It provides a way to audit the model's reasoning.

Meng: If I can jump in here, Tom. That interpretability piece is what gets me excited as an engineer. It's not just a black box saying "pneumonia." It's showing its work, which is the first step toward building trust and debugging failures.

Tom: Great point, Meng. And that leads to the final question: what does this all mean for the future of medical AI? Let's wrap this up.

Conclusion: Tom: So, we've spent some time with "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs." Let's pull it all together. What's the big takeaway?

Jane: The big takeaway is that eye-tracking data isn't just a way to see where someone looked. It's a way to capture *how* an expert thinks. By treating the gaze sequence as a form of reasoning supervision, this team has shown you can train a VLM to be both more accurate and more robust.

Tom: And that's a really elegant idea. It's not about adding more data, but about using the data you have in a smarter way. The gaze is a free, rich signal that's already being recorded, and this paper shows a concrete way to leverage it.

Jane: And the results speak for themselves. State-of-the-art on the in-domain test, and it generalizes better to unseen datasets. That's the kind of progress that could actually move the needle in a clinical setting.

Tom: It also opens up a whole new avenue of research. Could we use this for other medical imaging modalities, like CT or MRI? Could we use it to teach models to search for specific pathologies more effectively?

Jane: Absolutely. And it makes the AI's decision-making process more transparent. That's going to be critical for getting clinicians to trust and adopt these tools. It's not just a black box anymore; it's a model that can show you its reasoning, step by step.

Tom: Well said. We'll be keeping an eye on this line of work, pun intended. Thanks to everyone for tuning in. We'll be back soon with another paper.

Jane: See you next time, everyone.

Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao

University of Georgia · University of Texas, Arlington · Massachusetts General Hospital, Harvard Medical School · New Jersey Institute of Technology

cs.CV, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 63/100

Key concepts

Sequential Eye-Tracking
This refers to the specific order in which a radiologist's eyes move across an image when reading it. The paper argues this sequence represents the expert's actual reasoning process, moving beyond just where a person looked spatially.
Visual Reasoning Supervision
This is the method used to train vision-language models. Instead of just teaching the model what to see, this technique uses the time-ordered gaze data as supervision, teaching the model *how* to look for evidence step-by-step.
Gaze Tokens
These are four special placeholder tokens added to a language model's vocabulary. The model is trained to output these at the start of its answer, and then its internal states are used to predict which specific part of the image the radiologist was looking at.
Zero-Shot Generalization
This test measures how well a trained model performs on completely new datasets it has never seen before. The paper showed that the gaze-trained model performed best on three different chest X-ray datasets without any further fine-tuning.

Terminology

Summary

Summary

This paper introduces a method to make vision-language models (VLMs) for chest X-ray interpretation more visually grounded by using temporally ordered eye-gaze data from radiologists as supervision. The authors argue that Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. They propose using eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration.

The method is built on the observation that "despite multimodal inputs, however, many VLM pipelines still rely on text-only intermediate reasoning: the model first converts visual evidence into textual descriptions and then 'thinks' in language. This can be suboptimal for tasks that are inherently visual, where the most informative signals are not easily verbalized without loss. The authors note that gaze is not merely an attention map; it is a time-ordered record of how experts gather evidence, which is well aligned with the token-by-token computation of modern VLMs."

The approach uses the MIMIC-EYE dataset, which provides chest radiographs with synchronized eye-tracking signals and transcribed speech produced during diagnostic reading. The gaze signal is converted into attention heatmaps over the image to represent regions the radiologist focuses on, then discretize the heatmap into a fixed patch grid consistent with the VLM visual tokenizer. Gaze supervision is stored as a set of patch indices (top-k patches per gaze token), which is compact and directly compatible with a classification head over patch IDs.

The model architecture uses a pretrained VLM backbone (Qwen2.5-VL-7B-Instruct) with exactly four special placeholder tokens (denoted as ⟨st⟩1,..., ⟨st⟩4) at the beginning of the assistant response. The assistant is trained to output: ⟨st⟩1 ⟨st⟩2 ⟨st⟩3 ⟨st⟩4 Answer: 14 findings as yes/no. A Gaze projection head maps the hidden states at these four token positions to logits over P image patches using a linear projection. A separate 14-label classifier head is attached to the last token hidden state for multi-label prediction.

Training uses a two-stage optimization strategy. Stage 1 focuses on learning a consistent mapping between the four gaze tokens and gaze-selected patch indices using a cross-entropy loss over patch IDs, with masking for tokens lacking gaze targets. Stage 2 trains a 14-label classifier head while maintaining the constrained answer format using binary cross-entropy, optionally combined with a language modeling loss. The loss weight is set to λ = 0.7. Parameter-efficient fine-tuning with LoRA is used, keeping the backbone largely frozen.

Experiments on the MIMIC-EYE test set show that supervised fine-tuning (SFT) yields a large gain in AUROC (49.74 → 87.60). Adding gaze supervision as token-level patch supervision with original temporal order gives the best in-domain performance: Original-Gaze achieves the best in-domain performance, reaching 90.17 AUROC, outperforming Shuffled-Gaze (88.51) and Random-Gaze (86.45). The authors state this trend supports our hypothesis that gaze is not only a spatial prior: its sequential structure reflects expert evidence acquisition and aligns well with token-based VLM computation.

For zero-shot generalization on three external benchmarks (CheXpert 5×200, RSNA, SIIM-ACR), Our Original-Gaze is the best-performing method on every benchmark, achieving 62.45/61.73 (Acc/F1) on CheXpert 5×200, 77.61/53.73 on RSNA, and 64.07/61.89 on SIIM-ACR. Ablations show that both Random-Gaze and Shuffled-Gaze often outperform SFT, but preserving the original gaze order yields the most consistent and largest improvements, especially on the harder metrics such as F1.

The authors conclude that gaze-supervised token learning improves not only in-domain accuracy but also out-of-domain robustness, suggesting that ordered gaze encodes transferable visual evidence cues that benefit VLM-based medical classifiers. The contributions are summarized as: Gaze-guided reasoning supervision for radiology VLMs, State-of-the-art accuracy with clinician-friendly interpretability, and Stronger out-of-domain robustness.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system:

  • Implementation: Add a dedicated module that reserves 4 special placeholder tokens (1–4) at the start of the assistant response. These tokens are trained via a linear projection head to predict gaze-selected image patch indices (cross-entropy loss).

  • Benefit: The model learns to output intermediate visual reasoning steps that mimic radiologists' sequential evidence acquisition, rather than jumping directly to text-based conclusions.

  • Stage 1: Train LoRA adapters + gaze projection head using cross-entropy loss on patch indices, masking tokens with missing gaze targets. This teaches the model where to look and in what order.

  • Stage 2: Add a 14-label binary classifier head (BCE loss) while maintaining the fixed-format yes/no output. Combine with language modeling loss (λ=0.7 for classifier).

  • Benefit: The model internalizes temporal gaze patterns, not just spatial attention. This is critical because shuffled gaze (random order) underperforms original gaze by 1.66 AUROC on MIMIC-EYE.

  • Implementation: Enforce strict output format: 1 2 3 4 Answer: [14 findings yes/no]. Extract hidden states at these token positions for gaze supervision.

  • Benefit: Guarantees stable extraction of gaze-related representations, enabling reliable training and inference. Prevents format drift that would break the gaze-token alignment.

  • Implementation: Keep the Qwen2.5-VL-7B backbone frozen, train only LoRA adapters and the gaze projection head. Use 8×24GB A6000 GPUs.

  • Benefit: Reduces memory footprint by 90% while achieving state-of-the-art performance (90.17 AUROC on MIMIC-EYE), making it feasible for clinical deployment on standard hardware.

  • Implementation: Convert continuous gaze heatmaps into top-k patch indices per gaze token, aligned with the VLM's visual tokenizer grid. Handle missing gaze data (blinks, dropouts) by masking those tokens during loss computation.

  • Benefit: Compact representation (patch IDs) that is directly compatible with classification heads, and robust to noisy gaze signals.


  • AUROC: 90.17% (vs. 87.60% for SFT baseline, +2.57%)

  • Accuracy: 89.02% (vs. 86.03%, +2.99%)

  • F1-score: 87.61% (vs. 84.18%, +3.43%)

  • Key capability: The system now follows radiologist-like visual search patterns, revisiting suspicious regions in temporal order, leading to more accurate multi-label classification of 14 chest X-ray findings.

  • CheXpert 5×200: 62.45% accuracy, 61.73% F1 (best among all methods)

  • RSNA: 77.61% accuracy, 53.73% F1 (best among all methods)

  • SIIM-ACR: 64.07% accuracy, 61.89% F1 (best among all methods)

  • Key capability: The system generalizes to out-of-distribution data without fine-tuning, because it learns transferable evidence acquisition patterns rather than dataset-specific shortcuts. This is critical for deployment across different hospitals, imaging protocols, and patient populations.

  • Gaze-linked evidence: For each prediction, the system outputs 4 gaze token positions that indicate which image patches were attended to and in what order.

  • Clinical utility: Radiologists can audit the model's reasoning by reviewing the gaze sequence—if the model looked at the right regions in a plausible order, the prediction is more trustworthy. This supports retrospective review and error analysis.

  • Masking mechanism: When gaze data is unavailable (e.g., blinks, head motion, or non-eye-tracked datasets), the corresponding tokens are simply masked during training, and the model still works at inference time.

  • Benefit: The system degrades gracefully—it doesn't require eye-tracking at deployment, only during training.

  • Hardware: Runs on 8×24GB GPUs (e.g., RTX A6000), feasible for hospital IT infrastructure.

  • Inference speed: Fixed-format output (no free-form text generation) makes inference fast and deterministic.

  • Compatibility: Works with any VLM backbone (Qwen2.5-VL used here, but architecture is model-agnostic).

Aspect Baseline (SFT) Improved System Gain


Reasoning type Text-only intermediate steps Latent visual tokens with gaze order +2.57 AUROC

Attention mechanism Static spatial attention Temporal, ordered visual search +1.66 AUROC vs. shuffled gaze

Generalization Dataset-specific shortcuts Transferable evidence acquisition +6.85 F1 on RSNA

Interpretability Black-box Gaze-linked patch evidence Qualitative

The key innovation is treating gaze as a temporal supervision signal, not just a spatial prior. This aligns with how VLMs compute token-by-token, making the model think with gaze rather than look with gaze.

Abstract

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

Sources

Related papers