LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation

arXiv:2602.04617 · cs.CL · Submitted 2026-02-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation".

Jane: Layer-wise Expert-aligned Decoding (LEAD) is a novel framework designed to mitigate hallucinations in radiology report generation by inherently modifying the Large Vision-Language Model's decoding trajectory.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So to wrap up our discussion on "LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation," we’ve covered how this framework uses multiple visual experts injected at every layer of the decoder to steer the AI toward factual consistency, and how it balances generation loss with expert supervision.

Jane: I think the key takeaway here is that LEAD moves beyond simple external guidance by modifying the model's intrinsic decoding path itself, using a context-aware gated fusion mechanism to adaptively absorb visual information during text generation. This directly addresses the issue of image-ungrounded pathological details.

Lu: The implication for future research is clear: we should explore how these expert injection mechanisms can be generalized beyond radiology reports to other multimodal tasks where factual grounding against specific input modalities is required, pushing the boundaries of vision-language alignment.

Meng: For practical deployment, it suggests that focusing on the quality and diversity of those pathological features used to train the experts will be just as important as optimizing the model architecture itself for achieving reliable clinical performance.

Lalam: I think what this paper really points toward is a more trustworthy AI component in medicine; if we can achieve fidelity like this, it sets a higher bar for how we build generative models that interact with sensitive data.

Tom: It’s exciting to see how they managed to maintain linguistic coherence while simultaneously aligning the output with visual facts through these layer-wise adjustments. That approach is certainly something worth exploring further in the community.

Conclusion: Tom: So we’ve seen how LEAD uses visual expert signals at every layer to make sure those radiology reports are actually grounded in what the image shows, and now it's time to talk about what this whole thing means for AI in healthcare.

Jane: Exactly, Tom; the title itself tells us that this framework is all about aligning the language generation process with specific visual knowledge from experts. I think we should focus on how they achieved that alignment without just making the model slower or less fluent.

Lu: From a theoretical standpoint, I find it fascinating how they treat every decoder layer as an interactive consultation, essentially giving the large model real-time access to specialized visual expertise at each step of text creation. That's a very creative way to think about knowledge retrieval during generation.

Meng: Creatively is one word, Lu; I need to know if this means we can actually deploy something that meets clinical standards in a fast, reliable way without needing massive computational resources for every single inference step. Practicality is my main concern here.

Lalam: I see this as a significant step toward building AI that doesn't just predict the next plausible word, but one that respects the underlying reality of what's happening in a patient’s scan, which could fundamentally improve how we trust diagnostic tools overall.

Tom: That’s a big picture thought, Lalam; it moves us closer to truly reliable assistive technology instead of just fancy text generators. Jane, can you explain the core idea of this expert-guided decoding in plain language for our listeners?

Jane: Certainly, Tom; basically, imagine the AI writing a report and at every single moment it writes a sentence, a visual expert steps in to check if what it's saying matches what's actually visible in the X-ray. It’s like having an assistant constantly whispering factual checks into its ear while it types.

Lu: And that whisper isn't random; they use a sophisticated gating mechanism to decide how much weight that visual expert should have at any given layer, which is where the real innovation lies in making it context-aware.

Meng: That sounds complex to implement on a production system; what does this mean for the training pipeline? Do we need to painstakingly label every intermediate step with expert feedback? I'm wondering about the engineering overhead involved in setting up that entire visual expert module.

Lalam: The training strategy they used, combining the generation loss with explicit classification loss for those visual experts, seems like a smart way to ensure the model learns both to write well and to know when it needs visual confirmation. That dual supervision is key.

Tom: So we’ve seen that LEAD is this structured method of injecting specialized visual information layer by layer, and now we need to consider what this means for the future of AI applications in medical imaging. Where does this leave us?

Beijing Institute of Technology · Zhongguancun Academy

cs.CL

Submitted: 2026-02-04

Updated: 2026-10-01

Importance score: 83/100

The gist: Layer-wise Expert-aligned Decoding (LEAD) is a novel framework designed to mitigate hallucinations in radiology report generation by inherently modifying the Large Vision-Language Model's decoding

Key concepts

Visual Expert Module
This module uses separate classifiers for different pathology categories (like cardiomegaly) to extract specific visual features. These features are combined into a single 'expert embedding' that represents detailed medical findings from the image.
Layer-wise Expert-Aligned Decoding
This is the core process where expert signals are injected at every decoder layer. A dynamic gate checks if the current text context matches visual facts; if so, it opens to allow expert information into the model's hidden state, steering generation toward accuracy.
Context-aware Gated Fusion
This mechanism determines how much expert information to use at each step. It calculates a dynamic gate based on both the current text context and the visual signal. This soft selection switch ensures that relevant visual data is incorporated only when it aligns with what the model is currently writing.
Composite Loss Function
The training uses two losses: one for standard text generation (Cross-Entropy) and another for classifying pathologies using Binary Cross-Entropy. Balancing these losses forces the model to learn both fluent language and accurate visual grounding simultaneously.

Terminology

Summary

Layer-wise Expert-aligned Decoding (LEAD) is a novel framework designed to mitigate hallucinations in radiology report generation by inherently modifying the Large Vision-Language Model's decoding trajectory. This method addresses the problem of plausible yet image-ungrounded pathological details by injecting fine-grained visual expert signals into intermediate decoder representations at every inference step, dynamically rectifying decoding biases and steering the generation toward factual consistency.

The gist: LEAD proposes a novel framework that uses multi-label pathological features as visual expert signals to rectify the intrinsic decoding priors of LLM by adaptively injecting these signals into each decoder layer via a context-aware gated fusion mechanism.

How it works

The core idea of LEAD is to treat the LLM decoding process as an iterative consultation between a student language model and visual experts who provide fine-grained pathological features. This is achieved by designing a multiple experts module that extracts distinct pathological features, which are then integrated into each decoder layer via a gating mechanism. This layerwise architecture enables the LLM to consult expert features at every inference step via a learned gating function, thereby dynamically rectifying decoding biases and steering the generation toward factual consistency.

Visual Expert Module

To extract pathological visual features, the framework integrates a visual expert branch comprising pathology-specific classifiers corresponding one-to-one with target categories. Each expert is instantiated as a three-layer MLP binary classifier, supervised by pathological labels extracted from reports (e.g., cardiomegaly, pleural effusion). The intermediate features from each classifier are concatenated to construct a comprehensive expert embedding. This expert embedding is then incorporated into the LLM decoding process through a fusion mechanism.

Layer-wise Expert-Aligned Decoding Method

The LEAD method employs a context-aware gated fusion mechanism at each decoder layer to adaptively absorb expert information. Specifically, for the l-th layer, the injection intensity is determined by checking both the current context and the visual signal to compute a dynamic gate: g lt = σϕgate([h lt; e l]) (1), where [·; ·] represents concatenation. The expert signals are then injected via an interpolation connection: h'lt = (1 − g lt) ⊙ h lt + g lt ⊙ e l (2). This mechanism functions as a soft selection switch, ensuring that if the current hidden state correlates with text grounded in visual facts, the gate opens to incorporate relevant expert information, effectively steering intermediate representations toward a decoding trajectory aligned with visual facts.

Training Strategy

Training is optimized using a composite loss function L combining the generation loss and the pathological classification loss: L = Lgen + λLcls (5). The generation task is optimized via standard Cross-Entropy loss (Lgen) for next-token prediction, while the expert module is supervised using a multi-label Binary Cross-Entropy loss (Lcls) on the predictions of the expert classifiers. The total objective balances these terms: λ = 4 to achieve a balance of magnitude between the losses, enforcing explicit supervision on the visual experts to ensure injected signals are factually grounded. The model backbone is fine-tuned using Low-Rank Adaptation (LoRA) while freezing core parameters to prevent catastrophic forgetting, and the vision encoder and proposed modules are fully fine-tuned.

Key Contributions

The paper makes several key contributions:

  1. Proposing Layer-wise Expert-aligned Decoding, a novel framework that uses multi-label pathological features as visual expert signals to rectify the intrinsic decoding priors of LLM by strategically injecting these signals into each decoder layer.

  2. Designing a bidirectional integration mechanism that treats report generation as an interactive expert-guided process, employing a context-aware gated fusion mechanism at each decoding layer to adaptively integrate fine-grained expert signals.

  3. Demonstrating that this approach substantially improves clinical accuracy and effectively reduces hallucinations on both the CheXpert Plus and MIMIC-CXR datasets, showing that internal decoding guidance is critical as external constraints.

Results Summary

Experiments on the CheXpert Plus dataset show that LEAD yields effective improvements in clinical accuracy metrics while mitigating hallucinations. For instance, using Qwen3-8B, it achieves a clinical F1-score of 0.275, surpassing models like MambaXray-VL and R2GenCSR. Furthermore, the ablation study confirms the necessity of each component: replacing the dynamic gate with a direct Add operation leads to suboptimal performance, highlighting the critical role of the context-adaptive gated fusion in maximizing visual alignment without disrupting linguistic coherence. Qualitative analysis further demonstrates that LEAD accurately rectifies decoding biases by capturing fine-grained pathological details missed by other methods, such as detecting pleural effusions and pulmonary edema. The results confirm that even our lightweight 2B variant achieves an F1-score of 0.243, indicating effectiveness stems from the intrinsic alignment mechanism rather than parameter scaling.

Improvements for AI systems

As a fastidious researcher, I have analyzed the proposed Layer-wise Expert-aligned Decoding (LEAD) framework for Radiology Report Generation (RRG). The paper identifies a critical weakness in current Large Vision-Language Models (LVLMs)—their susceptibility to hallucinations due to intrinsic LLM decoding priors and poor cross-modal alignment.

Here are the specific, actionable improvements that can be implemented using the LEAD methodology, and what the resulting AI system will be capable of:


)

  1. No longer relying on external RAG or prompt-level guidance for hallucination mitigation; instead, implement an internal, dynamic mechanism to correct decoding biases.

  2. Integrate a multi-expert module (pathology classifiers) directly into the LLM’s intermediate representations at every decoder layer using a context-aware gated fusion mechanism to dynamically steer the generation trajectory.

  3. Employ confidence-aware feature aggregation to modulate expert signals, ensuring that only high-confidence pathological cues are injected into the decoding process, thereby suppressing noise from irrelevant experts.

)

The improved AI system will be capable of:

  1. Generating radiology reports with significantly higher factual consistency and clinical accuracy by directly aligning generated text with fine-grained visual evidence, rather than relying on external knowledge retrieval or static prompt modifications.

  2. Demonstrating robust hallucination suppression across different model scales (e.g., 2B, 4B, 8B parameters) and training configurations (frozen vs. fine-tuned), confirming the method's intrinsic robustness against LLM priors.

  3. Achieving state-of-the-art clinical efficacy metrics (e.g., higher F1 scores on CheXpert Plus) by selectively injecting pathological features that the backbone model might otherwise miss, leading to better detection of subtle findings like pleural effusions or pulmonary edema.

  4. Operating as a more reliable decision-support tool in clinical settings by producing reports that are not only fluent but are provably grounded in the specific visual pathology present in the medical image.

Sources

Related papers