Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

summary

Video file (mp4)

In short

The episode discusses a paper titled "Thinking with Gaze," which uses sequential eye-tracking data from radiologists to supervise medical vision-language models (VLMs). The team found that training models to follow a radiologist's visual search path improves accuracy and robustness on unseen datasets, offering better interpretability for clinical use.

Key concepts

Sequential Eye-Tracking
This refers to the specific order in which a radiologist's eyes move across an image when reading it. The paper argues this sequence represents the expert's actual reasoning process, moving beyond just where a person looked spatially.
Visual Reasoning Supervision
This is the method used to train vision-language models. Instead of just teaching the model what to see, this technique uses the time-ordered gaze data as supervision, teaching the model *how* to look for evidence step-by-step.
Gaze Tokens
These are four special placeholder tokens added to a language model's vocabulary. The model is trained to output these at the start of its answer, and then its internal states are used to predict which specific part of the image the radiologist was looking at.
Zero-Shot Generalization
This test measures how well a trained model performs on completely new datasets it has never seen before. The paper showed that the gaze-trained model performed best on three different chest X-ray datasets without any further fine-tuning.

Terminology used across episodes

This episode discusses

The paper

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs · Read on arXiv

Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao

University of Georgia · University of Texas, Arlington · Massachusetts General Hospital, Harvard Medical School · New Jersey Institute of Technology

Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs".

Jane: The paper was written by Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu et al. from University of Georgia and University of Texas, Arlington and Massachusetts General Hospital, Harvard Medical School and New Jersey Institute of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper with a title that just grabs you: "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs." Jane, I gotta say, that title is doing a lot of heavy lifting.

Jane: It really is, Tom. And it's from a team spanning the University of Georgia, UT Arlington, Harvard, and NJIT. The core idea is that when a radiologist reads a chest X-ray, they don't just stare at the whole image at once. Their eyes move in a specific sequence, hopping from one suspicious spot to another. That's the "gaze" part.

Tom: Right, and that sequence isn't random. It's the doctor's actual reasoning process made visible. The paper argues that most vision-language models, or VLMs, they take the image in, but then they "think" in text. They describe what they see and then reason about that description.

Jane: Exactly. And that's a bottleneck. Some visual information just doesn't translate well into words. So this team's idea is to use the radiologist's eye-tracking data as a kind of "thinking" supervision. They're teaching the model to follow the same visual path the expert did, step by step.

Tom: So instead of just saying "there's a shadow here," the model learns the *process* of finding that shadow. That's a pretty profound shift in how we think about training these models, isn't it?

Jane: It is. It's moving from teaching the model *what* to see, to teaching it *how* to look. And that's a much richer signal. The authors even call it "visual reasoning supervision." It's not just an attention map; it's a time-ordered record of evidence gathering.

Tom: And that temporal order is the key, right? We'll get into the details, but I'm already excited about the potential here. It's like giving the AI a mentor who shows it their work, not just the final answer.

Jane: And that mentor is a board-certified radiologist. The implications for training and for building trust in these systems are huge. Stick around, because we're going to dig into how they actually made this work.

Summary: Tom: So, Jane, we've got this great title. Now, what's the actual meat of "Thinking with Gaze"? Give me the quick version.

Jane: Okay, so they took a powerful VLM, specifically Qwen2 point 5-VL-7B, and they fine-tuned it on a dataset called MIMIC-EYE. That dataset has chest X-rays, the radiologist's report, and crucially, the synchronized eye-tracking data from when the radiologist was reading the image.

Tom: And that's the secret sauce. But how do you get a language model to "look" at something? It just predicts text tokens.

Jane: Right, so they did something clever. They added four special placeholder tokens to the model's vocabulary. Let's call them "gaze tokens." The model is trained to output these tokens at the start of its answer. Then, they take the internal hidden states at those token positions and use a small projection head to predict which patch of the image the radiologist was looking at.

Tom: So the model is literally being forced to predict the doctor's eye movements, in order, before it gives its final diagnosis. That's wild.

Jane: It is. And they do this in two stages. Stage one is all about learning this gaze prediction. Stage two is the actual classification task, where the model has to output a "yes" or "no" for fourteen different findings, like cardiomegaly or pneumonia. They also use LoRA, which is a clever trick to fine-tune the model without having to update all of its billions of parameters.

Tom: And the results? They compared their method, which they call "Original-Gaze," against a bunch of baselines. What happened?

Jane: On the in-domain MIMIC-EYE test set, their method hit an AUROC of ninety point one seven. That's a big jump from the supervised fine-tuning baseline, which was at eighty-seven point six zero. And it beat out other gaze-based methods, like just using a static heatmap.

Tom: So preserving the order of the gaze is important. It's not just about *where* they looked, but *when*.

Jane: Precisely. They even ran an ablation where they shuffled the gaze order, and performance dropped. That tells you the temporal sequence is a crucial part of the signal. It's not just a spatial prior; it's a temporal reasoning chain.

Tom: That's a really strong finding. It validates the whole premise of the paper. I'm curious to see how this holds up on data the model has never seen before. That's what we're going to talk about next.

Improvements: Tom: We're back with "Thinking with Gaze," and Jane just laid out the in-domain results. But the real test for any model like this is, does it work on new data? Does it generalize?

Jane: And that's exactly where this paper shines, Tom. They took their trained model and tested it zero-shot on three completely different chest X-ray datasets: CheXpert, RSNA, and SIIM-ACR. No fine-tuning on those datasets at all.

Tom: Zero-shot, meaning it has to figure it out from scratch. That's a tough test. And how did the gaze-trained model do?

Jane: It was the best performer on every single one of those benchmarks. For example, on the RSNA dataset, their "Original-Gaze" method got an F1 score of fifty-three point seven three, compared to forty-three point seven seven for the standard supervised fine-tuning. That's a massive improvement in a metric that's often hard to move.

Tom: Wow. So it's not just memorizing the MIMIC-EYE data. It's learning a transferable skill. It's learning *how* to look for evidence.

Jane: Exactly. And that's the key implication. The authors argue that by mimicking the radiologist's visual search pattern, the model is less likely to rely on dataset-specific shortcuts. It's learning a more robust, human-like strategy for gathering evidence.

Tom: So it's not just about being more accurate; it's about being more reliable. That's huge for clinical deployment, where you're going to see images from all sorts of different machines and settings.

Jane: Right. And it also has an interpretability angle. The model is producing a gaze sequence as part of its output. That means a clinician can look at that sequence and see, "Oh, the model looked at the left lung base first, then the cardiac silhouette." It provides a way to audit the model's reasoning.

Meng: If I can jump in here, Tom. That interpretability piece is what gets me excited as an engineer. It's not just a black box saying "pneumonia." It's showing its work, which is the first step toward building trust and debugging failures.

Tom: Great point, Meng. And that leads to the final question: what does this all mean for the future of medical AI? Let's wrap this up.

Conclusion: Tom: So, we've spent some time with "Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs." Let's pull it all together. What's the big takeaway?

Jane: The big takeaway is that eye-tracking data isn't just a way to see where someone looked. It's a way to capture *how* an expert thinks. By treating the gaze sequence as a form of reasoning supervision, this team has shown you can train a VLM to be both more accurate and more robust.

Tom: And that's a really elegant idea. It's not about adding more data, but about using the data you have in a smarter way. The gaze is a free, rich signal that's already being recorded, and this paper shows a concrete way to leverage it.

Jane: And the results speak for themselves. State-of-the-art on the in-domain test, and it generalizes better to unseen datasets. That's the kind of progress that could actually move the needle in a clinical setting.

Tom: It also opens up a whole new avenue of research. Could we use this for other medical imaging modalities, like CT or MRI? Could we use it to teach models to search for specific pathologies more effectively?

Jane: Absolutely. And it makes the AI's decision-making process more transparent. That's going to be critical for getting clinicians to trust and adopt these tools. It's not just a black box anymore; it's a model that can show you its reasoning, step by step.

Tom: Well said. We'll be keeping an eye on this line of work, pun intended. Thanks to everyone for tuning in. We'll be back soon with another paper.

Jane: See you next time, everyone.

More episodes

← Home