Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time".
Jane: Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: To continue our discussion on "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," let's focus specifically on the paper's thesis—what exactly is it arguing and why should we care about this approach? The core argument centers on the fact that multimodal large language models often suffer from hallucinations because textual tokens dominate generation, which undermines the value of perceptual inputs.
Jane: That’s precisely where the paper starts; they observe an inherent imbalance in how these models use different input modalities during inference, where visual or auditory tokens receive insufficient attention compared to text tokens. This imbalance forces the model to rely on textual language priors instead of grounded evidence, which is what causes the outputs to diverge from the provided perceptual inputs.
Lu: The paper tackles this by proposing LIME, a training-free framework designed specifically to bolster multimodal grounding by explicitly enhancing modality usage during decoding. It claims that this can be achieved without modifying any of the model's existing parameters, which is a big deal for deployment.
Meng: So the claim is that you can improve grounding and reduce hallucinations by intervening directly in the key-value representations during inference using this LIME method, rather than trying to fix the underlying architecture or training data? That sounds like it bypasses some of the heavy lifting involved in full model fine-tuning.
Lalam: If we can achieve better grounding this way, it suggests that MLLMs could become much more reliable tools for tasks that require understanding real-world visual or auditory context, which is crucial for building useful applications.
Tom: Exactly; the paper argues that because of this imbalance, and because they use LRP to quantify token contributions, they can define a relevance-based objective to promote increased reliance on perceptual inputs through inference-time updates. It matters because it suggests inference control is a viable path to better grounding.
Jane: So the main point is that LIME uses Layer-wise Relevance Propagation to figure out token contributions and then sets up an optimization goal that increases the explanatory contribution of multimodal tokens while constraining deviations from the original model distribution through KL divergence.
Lu: That mechanism, using M for modality and T for text, captures the overall contribution of each type of token, allowing them to steer the model's internal state toward utilizing modalities more effectively during output generation.
Meng: It sounds like a clever way to get targeted improvements without needing to overhaul the entire model structure or restart a massive training cycle from scratch. That makes it very appealing for incremental improvements in production environments.
Lalam: For Lalam, this means that if we deploy an MLLM on a complex visual task, this method could make its responses far more accurate because it won't just guess based on text patterns when there is clear visual evidence present.
Tom: Right; so to summarize the gist of "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," the paper proposes LIME as a training-free method that uses LRP to quantify token relevance and defines an objective to increase reliance on perceptual inputs during inference, addressing the textual dominance issue that causes hallucinations.
Jane: And it matters because it shows that controlling modality utilization at inference time is an effective strategy for improving multimodal grounding, which opens up new avenues for deploying more trustworthy AI.
Conclusion: Tom: So, wrapping up our conversation on "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," we have looked at how LIME addresses the problem of textual tokens dominating multimodal generation by using relevance propagation to steer inference updates toward stronger perceptual grounding. The authors are Joseph Keshet and Itai Allouche from Technion.
Jane: And the implication for us is that this work suggests a path forward where we don't necessarily need massive retraining efforts to fix hallucination issues; instead, we can introduce targeted relevance control during the inference phase itself. It shifts the focus from purely training-based fixes to inference-time behavioral modification.
Lu: From a research perspective, I think this is significant because it establishes that interpretability methods like LRP can be directly leveraged not just for post-hoc analysis but as active components in guiding model behavior during decoding, which is an interesting direction for how we study AI.
Meng: Practically speaking, the implication here is that if we want to deploy multimodal AI systems in high-stakes environments where accuracy matters, having a mechanism like LIME that can be toggled on for inference could dramatically increase our confidence in the outputs without needing continuous retraining cycles.
Lalam: For Lalam, this means that future versions of the AI I represent will be inherently more trustworthy when interacting with visual or audio data because they’ll have an internal mechanism to prioritize what's actually present in the input rather than just relying on textual probability.
Tom: That seems to be the big picture, Jane; it’s about adding a layer of control at inference time that specifically addresses the modality imbalance that causes these errors, making grounded multimodal AI more robust. We'll keep an eye on how this LIME method plays out in real-world benchmarks moving forward.
Technion
cs.LG, cs.CV, eess.AS
Submitted: 2026-05-03
Updated: 2026-09-28
Code: https://github.com/ItaiAllouche/lime
Importance score: 69/100
The gist: Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs.
Key concepts
- Hallucinations in MLLMs
- These are errors where the model generates factually incorrect or nonsensical information. In multimodal models, this often happens because the model relies too heavily on text tokens during generation instead of properly grounding its answers in visual or audio inputs, leading to inaccurate outputs.
- Layer-wise Relevance Propagation (LRP)
- LRP is a technique used to trace the contribution of each input token to the final output. In this paper, it was used first to prove that text tokens receive higher relevance scores than modality tokens during inference, confirming the imbalance causing hallucinations.
- Learning Inference-time Modality Enhancement (LIME)
- LIME is a training-free method that intervenes directly into the model's keyvalue representations during decoding. It uses a relevance-based objective to optimize these updates, forcing the model to pay more attention to perceptual inputs while staying close to its original knowledge.
Terminology
Summary
Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs. The proposed framework addresses this by introducing Learning Inference-time Modality Enhancement (LIME), a training-free method that explicitly enhances modality usage during decoding to bolster grounding without modifying model parameters.
The Gist
LIME leverages Layer-wise Relevance Propagation (LRP) to quantify token-level contributions and defines a relevance-based objective that promotes increased reliance on perceptual inputs through inference-time updates to the model’s keyvalue representations.
Motivation and Analysis of Hallucinations
The paper first establishes that multimodal hallucinations are closely related to an imbalance in how models utilize different input modalities during inference, where text tokens dominate the generation process, while visual or auditory tokens receive insufficient attention.
This is confirmed by applying Layer-wise Relevance Propagation (LRP), which uncovers a systematic imbalance: textual tokens are consistently assigned higher relevance than modality tokens, even in tasks that critically depend on perceptual input.
This pattern provides direct evidence that hallucinations are associated with the under-utilization of modality-specific information during inference.
The LIME Framework and Optimization
LIME is designed as a training-free framework to intervene directly in the model’s key-value (KV) representations. The core mechanism involves optimizing additive updates, denoted by ∆ = ∆K, ∆V, at each decoding step. To control this optimization, LIME employs a relevance-based objective inspired by Noise Contrastive Estimation (NCE) and contrastive learning principles:
- The relevance objective is defined as:
ΦM = Σi∈M Φi and ΦT = Σj∈T Φj, capturing the overall contribution of modality and textual tokens.
- The loss function to be minimized is Lrel(∆) + λLKL(∆), where Lrel(∆) increases the explanatory contribution of multimodal tokens while suppressing excessive reliance on textual context, and LKL(∆) constrains deviations from the pretrained model distribution by calculating DKL with respect to the next-token distribution pθ.
Evaluation and Results
The effectiveness of LIME is evaluated across multiple multimodal benchmarks in both vision and audio domains. The results demonstrate consistent reductions in hallucination and improved grounding while preserving generation quality.
Specific analyses show that LIME increases modality contribution and produces more localized and semantically aligned relevance patterns,
leading to improvements in both spatial grounding (e.g., POPE benchmark) and modality reliance across vision, audio, CHAIR, and AIR-Bench models.
Implementation Details and Trade-offs
The method operates by performing inference-time learning
through optimizable KV updates (∆KV),
which are computed independently at each decoding step and then discarded. The trade-off between performance and stability is managed by the KL regularization weight λ, where smaller values of λ enable stronger updates to the model’s internal representations, while larger values constrain the optimization and reduce its effect.
Computational overhead is noted, showing that LIME introduces increased latency compared to standard autoregressive decoding,
but this is deemed practical for settings where improved multimodal grounding is critical. Ablation studies confirm that jointly modifying both keys and values (∆KV) consistently achieves the best performance.
Conclusion
LIME successfully mitigates multimodal hallucinations by controlling modality utilization at inference time, proving that inference-time control of modality contributions is an effective strategy for improving the reliability of MLLMs.
The findings suggest that this approach promotes more effective utilization of modality information during inference, leading to improved multimodal grounding.
The gist
LIME leverages Layer-wise Relevance Propagation (LRP) to quantify token-level contributions and defines a relevance-based objective that promotes increased reliance on perceptual inputs through inference-time updates to the model’s keyvalue representations.
How it works
-
The framework first employs LRP, using Attention-Aware Layer-wise Relevance Propagation (AttnLRP), to compute token-level relevance scores over input tokens X, resulting in total relevance quantities ΦM and ΦT for modality and textual tokens.
-
The goal is to find the additive updates ∆ = ∆K, ∆V at each decoding step that shift relevance toward modality tokens while remaining close to the original model distribution.
-
This is achieved by minimizing a composite loss: Lrel(∆) + λLKL(∆). The Lrel objective uses a temperature-scaled softmax over relevance scores to increase the explanatory contribution of multimodal tokens, while the LKL term regularizes modifications using KL-divergence with respect to the reference mode pθ.
Relevance Propagation Mechanics
The paper details specialized propagation rules for transformer architectures, including:
-
The standard LRP-z rule for linear transformations, which is adapted for transformer layers.
Improvements for AI systems
Here are specific improvements for AI systems based on the LIME framework proposed in this paper:
-
Improve Multimodal Grounding in Hallucination-Prone MLLMs (Vision/Audio):
-
Enhance Perceptual Fidelity and Object Verification: The improved system will significantly reduce the generation of hallucinations (e.g., describing non-existent objects or events) by explicitly enforcing reliance on modality tokens (visual features or audio segments) during the decoding process, rather than relying solely on textual priors.
-
Improve Spatial and Temporal Precision in Multimodal Understanding: By optimizing relevance propagation based on Layer-wise Relevance Propagation (LRP), the system will produce more localized and semantically aligned relevance patterns. This allows the model to focus its attention precisely on ground-truth regions (e.g., specific bounding boxes in images or exact temporal segments in audio), leading to more accurate answers for VQA, captioning, and sound event detection tasks.
-
Increase Modality Reliance for Complex Reasoning: The framework will systematically increase the proportion of relevance attributed to modality tokens relative to textual inputs. This forces the model to utilize perceptual evidence as a primary source of truth during inference, which is crucial for complex reasoning tasks that depend on grounding in sensory input rather than just linguistic knowledge.
-
Develop Robust Training-Free Mitigation Strategies: The LIME framework provides a training-free method for mitigating hallucinations by intervening directly in the Key/Value (KV) representations at inference time. This allows developers to enhance multimodal grounding without requiring costly model fine-tuning or additional labeled data, making it a practical solution for deploying high-reliability MLLMs in latency-sensitive environments.
-
Achieve Improved Performance Across Diverse Modalities: The system can be applied consistently across both vision and audio domains (e.g., using LLaVA/Qwen models with images and Whisper/Audio encoders). This generalization suggests the framework can effectively address modality imbalance issues across different sensory inputs, leading to consistent reductions in hallucination rates on benchmarks like POPE, CHAIR, and AIR-Bench.
-
Fine-tune Optimization Trade-offs: By incorporating a KL-divergence regularizer alongside the relevance objective, the system allows for fine control over the trade-off between promoting modality utilization (reducing hallucinations) and preserving linguistic coherence/decoding stability of the original pretraining (preventing catastrophic forgetting or instability).
Sources
- MIRAGE: The Illusion of Visual Understanding
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- Qwen2-Audio Technical Report
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
- Adam: A Method for Stochastic Optimization
- A Survey on Hallucination in Large Vision-Language Models
- Representation Learning with Contrastive Predictive Coding
- V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
- Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks