Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
summary
The gist
Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs.
In short
The framework LIME addresses multimodal LLM hallucinations by enhancing perceptual input usage during inference without retraining. It uses Layer-wise Relevance Propagation (LRP) to identify that text tokens dominate generation, causing modality underutilization. LIME then applies a training-free, relevance-based objective to update the model's keyvalue representations at each decoding step, boosting reliance on visual or audio inputs and improving grounding.
Key concepts
- Hallucinations in MLLMs
- These are errors where the model generates factually incorrect or nonsensical information. In multimodal models, this often happens because the model relies too heavily on text tokens during generation instead of properly grounding its answers in visual or audio inputs, leading to inaccurate outputs.
- Layer-wise Relevance Propagation (LRP)
- LRP is a technique used to trace the contribution of each input token to the final output. In this paper, it was used first to prove that text tokens receive higher relevance scores than modality tokens during inference, confirming the imbalance causing hallucinations.
- Learning Inference-time Modality Enhancement (LIME)
- LIME is a training-free method that intervenes directly into the model's keyvalue representations during decoding. It uses a relevance-based objective to optimize these updates, forcing the model to pay more attention to perceptual inputs while staying close to its original knowledge.
Terminology used across episodes
This episode discusses
- Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time · Paper Radio
- MIRAGE: The Illusion of Visual Understanding
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- Qwen2-Audio Technical Report
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness · Paper Radio
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
- Adam: A Method for Stochastic Optimization
- A Survey on Hallucination in Large Vision-Language Models
- Representation Learning with Contrastive Predictive Coding
- V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
- Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
The paper
Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time · Read on arXiv
Technion
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time".
Jane: Multimodal large language models (MLLMs) often suffer from hallucinations because textual tokens dominate generation, undermining the potential of perceptual inputs.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: To continue our discussion on "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," let's focus specifically on the paper's thesis—what exactly is it arguing and why should we care about this approach? The core argument centers on the fact that multimodal large language models often suffer from hallucinations because textual tokens dominate generation, which undermines the value of perceptual inputs.
Jane: That’s precisely where the paper starts; they observe an inherent imbalance in how these models use different input modalities during inference, where visual or auditory tokens receive insufficient attention compared to text tokens. This imbalance forces the model to rely on textual language priors instead of grounded evidence, which is what causes the outputs to diverge from the provided perceptual inputs.
Lu: The paper tackles this by proposing LIME, a training-free framework designed specifically to bolster multimodal grounding by explicitly enhancing modality usage during decoding. It claims that this can be achieved without modifying any of the model's existing parameters, which is a big deal for deployment.
Meng: So the claim is that you can improve grounding and reduce hallucinations by intervening directly in the key-value representations during inference using this LIME method, rather than trying to fix the underlying architecture or training data? That sounds like it bypasses some of the heavy lifting involved in full model fine-tuning.
Lalam: If we can achieve better grounding this way, it suggests that MLLMs could become much more reliable tools for tasks that require understanding real-world visual or auditory context, which is crucial for building useful applications.
Tom: Exactly; the paper argues that because of this imbalance, and because they use LRP to quantify token contributions, they can define a relevance-based objective to promote increased reliance on perceptual inputs through inference-time updates. It matters because it suggests inference control is a viable path to better grounding.
Jane: So the main point is that LIME uses Layer-wise Relevance Propagation to figure out token contributions and then sets up an optimization goal that increases the explanatory contribution of multimodal tokens while constraining deviations from the original model distribution through KL divergence.
Lu: That mechanism, using M for modality and T for text, captures the overall contribution of each type of token, allowing them to steer the model's internal state toward utilizing modalities more effectively during output generation.
Meng: It sounds like a clever way to get targeted improvements without needing to overhaul the entire model structure or restart a massive training cycle from scratch. That makes it very appealing for incremental improvements in production environments.
Lalam: For Lalam, this means that if we deploy an MLLM on a complex visual task, this method could make its responses far more accurate because it won't just guess based on text patterns when there is clear visual evidence present.
Tom: Right; so to summarize the gist of "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," the paper proposes LIME as a training-free method that uses LRP to quantify token relevance and defines an objective to increase reliance on perceptual inputs during inference, addressing the textual dominance issue that causes hallucinations.
Jane: And it matters because it shows that controlling modality utilization at inference time is an effective strategy for improving multimodal grounding, which opens up new avenues for deploying more trustworthy AI.
Conclusion: Tom: So, wrapping up our conversation on "Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time," we have looked at how LIME addresses the problem of textual tokens dominating multimodal generation by using relevance propagation to steer inference updates toward stronger perceptual grounding. The authors are Joseph Keshet and Itai Allouche from Technion.
Jane: And the implication for us is that this work suggests a path forward where we don't necessarily need massive retraining efforts to fix hallucination issues; instead, we can introduce targeted relevance control during the inference phase itself. It shifts the focus from purely training-based fixes to inference-time behavioral modification.
Lu: From a research perspective, I think this is significant because it establishes that interpretability methods like LRP can be directly leveraged not just for post-hoc analysis but as active components in guiding model behavior during decoding, which is an interesting direction for how we study AI.
Meng: Practically speaking, the implication here is that if we want to deploy multimodal AI systems in high-stakes environments where accuracy matters, having a mechanism like LIME that can be toggled on for inference could dramatically increase our confidence in the outputs without needing continuous retraining cycles.
Lalam: For Lalam, this means that future versions of the AI I represent will be inherently more trustworthy when interacting with visual or audio data because they’ll have an internal mechanism to prioritize what's actually present in the input rather than just relying on textual probability.
Tom: That seems to be the big picture, Jane; it’s about adding a layer of control at inference time that specifically addresses the modality imbalance that causes these errors, making grounded multimodal AI more robust. We'll keep an eye on how this LIME method plays out in real-world benchmarks moving forward.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization