Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs".
Tom: The gist The method introduces FIRE-MPO, a fine-grained,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we're talking about this paper now, "Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs." It seems like they’ve really zeroed in on how to make these vision language models actually pay attention to what matters in a medical context.
Jane: Right, it’s all about the limitations of current methods when you try to align these models for medical tasks. The authors point out that existing tools like DPO aren't handling the nuance well because they treat every part of an answer equally, even if some parts are crucial for diagnosis.
Lu: They focus on three specific problems with current alignment techniques in medicine: first, sequence-level rewards treat critical tokens just like filler text, which is a big issue when you need precision in a clinical setting.
Meng: And second, relying on old supervised fine-tuning references introduces a kind of shift in how the model learns, pushing it toward surface-level style instead of actual medical accuracy.
Lalam: Plus, they note that current alignment objectives just don't have this explicit visual grounding needed to make sure the model notices subtle but diagnostically important features in an image.
Tom: Exactly, so this paper is proposing a new framework to fix those issues by introducing a more targeted approach. We need to see how they actually build that fix in the next part of our discussion.
Jane: We'll talk about what exactly they are trying to do with this fine-grained preference optimization method and why it matters for medical AI right now.
The paper's summary: Tom: Okay, so the core idea is that they introduce FIRE-MPO, which is this new framework designed to handle those coarse reward signals by making the alignment happen at a token level. They’re using something called a bidirectional token-wise KL regularizer along with a visual-contrastive grounding objective.
Jane: That sounds technical, but the practical aim is clear: they want the model to be better at focusing on clinically significant visual regions when generating an answer. They also want to stop it from giving confident answers based on corrupted visual data.
Lu: The method structures the ranking loss around dual constraints, setting up pairs where they guide the model toward medically accurate answers supported by valid visuals, and also penalize it for hallucinating when the underlying image evidence is actually bad.
Meng: It sounds like they are building a system that forces the AI to check its claims against specific visual evidence rather than just following a general preference pattern. That has some serious implications for reliability in diagnostics.
Lalam: The paper talks about constructing these preference pairs by editing only the clinically incorrect parts of an output while keeping the rest of the response structure intact, which is a clever way to generate training data that's actually useful for refinement.
Tom: And they’re using this token-level focus to directly address how standard DPO treats every word in a response uniformly during the optimization process. This seems like a direct hit on the problem we talked about earlier.
Jane: It really boils down to making sure that when the AI is learning, it pays attention to *where* in an image the relevant information is, not just *that* there is something relevant somewhere in the whole picture.
The paper's improvements: Tom: Let’s get into the specific changes they made. First, they replaced those unidirectional constraints with a bidirectional token-wise KL regularizer. This allows them to balance keeping the model’s general knowledge while learning new medical domain specifics at the same time.
Jane: That balance is key, because if you only push one direction, you might lose old knowledge or fail to learn the new specific facts they are trying to teach it.
Lu: They also introduced a visual-contrastive grounding objective which pairs clean images with versions where lesions are corrupted to train the model specifically on distinguishing between what's supported and what isn't.
Meng: That part about contrasting clean images with lesion-corrupted ones sounds like a very direct way to teach the AI how to avoid making predictions when it can’t find visual support for a specific finding.
Lalam: They also refined how they construct those preference pairs, focusing on editing only the incorrect spans in an output rather than replacing the whole response with a completely different answer. That keeps the original structure of the model's output intact while fixing the medical errors.
Tom: So we have token-level rewards, bidirectional regularization for coverage, and explicit visual grounding objectives. That really shows how they’ve tried to build a more precise alignment process than what was there before.
Jane: The result of all this fine-tuning is that the model should show stronger visual grounding across different layers when it's performing medical tasks compared to the models they tested against.
Conclusion: Tom: So, wrapping up, these findings from "Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs" show a path toward better alignment for medical vision language models by focusing on fine-grained control at the token level. They’ve shown that this approach consistently outperforms standard methods on benchmarks like VQARAD and SLAKE.
Jane: The big picture here is that we’re moving away from these coarse sequence rewards and moving toward a system where the AI is genuinely tethered to localized pathological features in images for its responses. That should make the output more reliable when used in actual clinical settings.
Lu: It suggests that by making the alignment process more granular—token-by-token—we can capture the subtle visual cues that are essential for complex medical reasoning, which is a major step forward for how we train these models.
Meng: From an engineering standpoint, this means we have a clearer target for what to measure in our training data and how to build those grounding objectives to ensure the model doesn't just learn surface patterns but truly understands the underlying pathology.
Lalam: It’s really about making sure that when this AI is used for medical tasks, it’s not just producing fluent text, but text that is precisely grounded in the visual evidence it was trained on.
Tom: That’s the essence of what they did with FIRE-MPO. We've seen how fine-grained preference optimization can lead to more trustworthy AI in medicine.
York University · University of British Columbia · Vector Institute · Queen’s University
cs.CV, cs.AI
Submitted: 2026-06-10
Updated: 2026-10-08
Importance score: 90/100
The gist: The gist The method introduces FIRE-MPO, a fine-grained, on-policy alignment framework designed to address the unique challenges of medical Vision-Language Models by utilizing a bidirectional
Key concepts
- Sequence-level Reward Signals
- Current methods treat every word in an answer equally during training. This means important clinical details are given the same weight as generic filler phrases like 'the image shows.' This fails to prioritize medically critical tokens, leading to less accurate specialized responses.
- Bidirectional Token-wise KL Regularizer
- This technique uses two types of constraints (forward and reverse) applied at the individual token level. It balances preserving the model's existing broad knowledge with learning new, specific medical domain knowledge without losing important details during optimization.
- Visual-Contrastive Grounding Objective
- This objective trains the model to distinguish between responses supported by clear visual evidence and those that are visually unsupported or corrupted. This forces the AI to only make predictions based on localized, verifiable pathological features in medical images.
Terminology
Summary
The gist The method introduces FIRE-MPO, a fine-grained, on-policy alignment framework designed to address the unique challenges of medical Vision-Language Models by utilizing a bidirectional token-wise KL regularizer and a visual-contrastive grounding objective.
The limitations of existing methods
Existing post-training alignment approaches like Direct Preference Optimization (DPO) face three critical limitations in the medical domain (Page 1). First, sequence-level reward signals treat clinically critical tokens identically to generic filler text (Page 2). Second, reliance on static supervised fine-tuning references introduces an off-policy distribution shift, steering optimization toward stylistic artifacts over clinical correctness (Page 2). Third, existing alignment objectives lack explicit visual grounding constraints, leaving models insensitive to subtle yet diagnostically decisive pathological features (Page 2). Standard DPO treats all response tokens uniformly during optimization, meaning diagnostically important spans contribute no more to the training signal than generic filler phrases such as “the image shows” or “there is evidence of” (Page 1).
The FIRE-MPO framework components
FIRE-MPO addresses these limitations by introducing three main modifications to the ranking loss and the token-wise KL regularizer (Page 3).
-
Ranking Loss: The objective is structured around a dual-constraint, defining preference pairs (q, v, y+) ≻ (q, v, y−) to align πθ toward medically accurate answers when supported by valid visual evidence v (Page 3). It also introduces pairs (q, v′, y−) ≻ (q, v′, y+) to compel the model to refrain from providing a definitive response when the underlying visual evidence is corrupted (v′), thereby penalizing hallucinations that are not grounded in localized visual evidence (Page 3).
-
KL Regularizer: To balance coverage preservation and tail suppression, FIRE-MPO replaces unidirectional constraints with a bidirectional, tokenwise KL regularizer (Page 4). This involves applying both forward and reverse KL divergence constraints at the token level to allow for a balance between preserving the base model’s rich prior knowledge and learning new specialized domain knowledge (Page 6).
-
Preference Pair Construction: The framework constructs on-policy preference pairs directly from model-generated outputs by carefully editing only the clinically incorrect spans while preserving the model’s original linguistic structure, rather than replacing the entire response with a stylistically different reference answer (Page 2).
Key design choices and their rationale
The method incorporates several specific design choices to enhance medical alignment (Page 3). The use of a bidirectional token-wise KL regularizer is employed because it allows for balances coverage preservation and tail suppression by simultaneously applying both forward and reverse KL divergence constraints at the token level (Page 6). This balanced regularization ensures the policy achieves rigorous clinical accuracy without sacrificing the interpretive variety that is essential for complex medical reasoning (Page 4). Furthermore, the visual-contrastive grounding objective pairs clean images with lesion-corrupted counterparts to train the model to distinguish between clinically supported and visually unsupported responses, explicitly penalizing hallucinated predictions that are not grounded in localized visual evidence (Page 2).
Experimental validation and results
Experiments across multiple medical imaging benchmarks demonstrate that FIRE-MPO consistently outperforms both standard DPO and RRPO, achieving an average relative improvement of 10.24% across two state-of-the-art LVLMs (Page 2). Quantitative performance on HuatuoGPT-Vision-7B showed FIRE-MPO achieving the highest overall average score of 69.71%, improving upon the base model (61.21%) by 8.50 points, yielding a relative gain of +13.89% (Page 7). Qualitative analysis in deeper layers shows that FIRE-MPO demonstrates stronger visual grounding and more controlled distributional alignment across both the specialized HuatuoGPT-Vision-7B and the general-purpose Qwen3-VL-4B-Instruct backbones (Page 8).
Conclusion
FIRE-MPO is a fine-grained, onpolicy alignment framework that addresses the unique challenges of medical Vision-Language Models by moving beyond coarse sequence-level rewards, introducing a bidirectional tokenwise KL regularizer, and ensuring responses are authentically tethered to localized pathological features (Page 6). Extensive evaluations across VQARAD, SLAKE, and IU-Xray demonstrate that FIRE-MPO consistently outperforms state-of-the-art alignment methods (Page 6). The broader impact of this work lies in its potential to enhance the reliability and factual consistency of AI-driven medical diagnostics (Page 15).
The paper is titled Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs (Page 1). REFERENCES [1] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631
The gist The method introduces FIRE-MPO, a fine-grained, on-policy alignment framework designed to address the unique challenges of medical Vision-Language Models by utilizing a bidirectional token-wise KL regularizer and a visual-contrastive grounding objective.
How it works
Existing post-training alignment approaches like Direct Preference Optimization (DPO) face three critical limitations in the medical domain (Page 1). First, sequence-level reward signals treat clinically critical tokens identically to generic filler text (Page 2). Second, reliance on static supervised fine-tuning references introduces an off-policy distribution shift, steering optimization toward stylistic artifacts over clinical correctness (Page 2). Third, existing alignment objectives lack explicit visual grounding constraints, leaving models insensitive to subtle yet diagnostically decisive pathological features (Page 2). Standard DPO treats all response tokens uniformly during optimization, meaning diagnostically important spans contribute no more to the training signal than generic filler phrases such as “the image shows” or “there is evidence of” (Page 1).
Improvements for AI systems
- Bold header: Fine-grained token-level reward assignment
The system can correct only clinically erroneous spans while preserving the original linguistic style
by introducing a Finegrained Regularized Medical Preference Optimization (FIRE-MPO) objective, a fine-grained preference optimization objective that operates at the token level.
- Bold header: Mitigated stylistic reward hacking
The improved system will prevent models from optimizing for surface patterns by constructing on-policy preference pairs by carefully edit[ing] only the clinically incorrect spans and preserve[ing] the model’s original linguistic structure, rather than replacing the entire response with a stylistically different reference answer.
- Bold header: Explicit visual evidence grounding
The system will be insensitive to subtle features by introducing a visual-contrastive grounding objective that pairs clean images with lesion-corrupted counterparts,
which explicitly penalizes hallucinated predictions that are not grounded in localized visual evidence.
- Bold header: Balanced distributional control
By replacing unidirectional constraints with a bidirectional, tokenwise KL regularizer, the model can achieve balances coverage preservation and tail suppression by simultaneously applying both forward and reverse KL divergence constraints at the token level.
- Bold header: Enhanced visual grounding sensitivity
The system will demonstrate stronger visual grounding
across layers compared to baselines, as it is enabled by the fine-grained image preference construction pipeline,
progressively shifting attention toward diagnostically relevant localization and attribute tokens.
Sources
- Qwen3-VL Technical Report
- GPT-4 Technical Report
- OpenAI GPT-5 System Card
- DeepSeek-V3 Technical Report
- A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
- MedGemma Technical Report
- Capabilities of Gemini Models in Medicine
- Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
- MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
- Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
- Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
- 3D-Properties: Identifying Challenges in DPO and Charting a Path Forward
- Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
- Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
- Token-level Direct Preference Optimization
- TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
- Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
- Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models