MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models".
Jane: The paper was written by the authors from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: Okay, so we established that cross-modal hallucination is a serious problem when inputs are mixed, like images and text. Jane, if you could summarize what this paper suggests the core mechanism of MAD is? How does it fix the mixing problem?
Jane: The summary really focuses on this concept called 'Adaptive Decoding.' Instead of letting the model just generate tokens blindly after seeing all the modalities, MAD actively adapts its decoding process based on which modality seems most reliable for that specific part of the output.
Lu: Right, it’s not just about giving more attention to one modality; it’s about dynamically adjusting the decoding strategy itself. Think of it like having multiple filters that activate selectively depending on whether the input is primarily visual or linguistic, and which one provides better grounding for the current step.
Meng: So, if I understand this correctly, MAD isn't just adding another layer; it's modifying the *decoding* process itself to be more context-aware across modalities. That’s a fundamental change in how the output sequence is structured, which is much harder than just training better encoders.
Lalam: Because of that adaptability, the implications are that we can build systems that don't just sound plausible, but are provably grounded in the input data across all formats—be it video, text, or audio. That level of grounding restores confidence in AI understanding.
Improvements: Tom: It sounds like MAD is a big step up from previous methods that just tried to give the model more training data or bigger weights. Jane, what kind of specific improvements does the paper highlight? How does it improve upon the status quo?
Jane: What I gather is that MAD tackles this problem by making the decoding process itself conditional on the modalities available. It moves beyond simple fusion and into a sophisticated way of prioritizing information during generation, which was lacking before.
Lu: The improvement really lies in how it models modality interactions *during* inference time, not just training time. It suggests a more modular way of integrating cross-modal knowledge that prevents one type of input from corrupting the output generated from another.
Meng: From an implementation standpoint, this modularity sounds promising because it suggests we might be able to swap out components—say, if the vision encoder is updated, the core decoding logic adapts without requiring a complete overhaul of the entire system. That makes deployment much more manageable.
Lalam: This ability to dynamically improve reliability based on diverse inputs means that AI can assist in fields where ambiguity is common, like interpreting complex historical records or analyzing mixed-media art installations, grounding the interpretation in tangible evidence.
Conclusion: Tom: We've covered a lot of ground today regarding "MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models." Jane, if we had to wrap up the main implication for the average listener, what's the biggest takeaway?
Jane: The big idea is moving AI from being a creative storyteller that sometimes makes things up, to being an extremely reliable collaborator that always knows where its information comes from and how to use it correctly.
Lu: I think the most exciting long-term implication is that this architecture could unlock entirely new forms of scientific discovery by allowing AI to synthesize information across disciplines with unprecedented fidelity.
Meng: Honestly, if this level of reliability scales up, it changes the risk profile for implementing AI in critical infrastructure—we can finally start trusting it with mission-critical tasks.
Lalam: Ultimately, improving the reliability of multimodal AI isn't just a technical fix; it's a cultural enabler. It allows us to trust AI to augment our understanding of reality without constant skepticism about its factual basis.
Tom: Well, Jane, this has been an incredibly insightful discussion. Thank you to Lu, Meng, and Lalam for sharing your expertise with us today. We are definitely going to be following the developments surrounding "MAD: Modality-Adaptive Decoding
Conclusion: Tom: So, if we wrap all this up, what we're really looking at with "MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models" is a massive step toward trust.
Jane: Exactly, Tom. For so long, these multimodal AI systems have been fantastic at understanding the *parts*—the video, the sound—but they've often gotten tripped up by making things up when those parts didn't line up perfectly.
Meng: That’s the core problem that trips people up in real-world applications, isn't it? If you’re building something for safety or complex diagnostics, you can’t have the AI confidently hallucinating details that aren't actually present in the input.
Lu: But think about what this capability unlocks for scientific discovery! If an AI can reliably tell us, "The sound matches the action," rather than just guessing a correlation, we could revolutionize how we analyze raw sensor data from everything from deep-sea exploration to planetary missions.
Lalam: I agree with Lu; it fundamentally changes our trust model for technology. Knowing that the system is grounded in verifiable reality—the actual input—means that AI can become an even better partner in human culture, helping us preserve accuracy across all forms of media.
Jane: It really brings the concept of "grounding" into sharp focus, doesn't it? It makes the entire system more reliable because it’s constantly checking its own assumptions against multiple sources.
Tom: And that reliability is huge. Considering how much we rely on multimodal AI for everything from media analysis to medical imaging, this reduction in hallucination risk is absolutely critical for mainstream adoption.
Meng: Practically speaking, this means the engineering challenge shifts from just making bigger models to making smarter, more disciplined decoding processes. That's where the real money and progress are going to be.
Lu: It changes the entire paradigm of multimodal understanding; it’s moving us beyond mere pattern matching and into true contextual reasoning that respects physical laws and sensory input limitations.
Tom: You know, Jane, when we think about how much people interact with media every day, the ability for AI to be this disciplined is nothing short of revolutionary.
Jane: It makes me genuinely excited about what the next generation of multimodal tools will look like now that we have this safeguard in place.
Lalam: The fidelity it brings to understanding helps us appreciate human creativity and truth even more when we know the technology isn't pulling our strings with fabricated details.
Tom: Alright, team, that really wraps up our deep dive into "MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models." We gotta take a quick break, but when we come back, we’re going to be talking about some absolutely fascinating new developments in efficient AI architecture.
cs.AI
Submitted: 2026-01-29
Updated: 2026-09-11
Comments: Accepted to CVPR 2026
Code: https://github.com/top-yun/MAD
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 85/100
The gist: This paper presents Modality-Adaptive Decoding (MAD), a training-free method designed to mitigate "cross-modal hallucinations" in Multimodal Large Language Models (MLLMs).
Key concepts
- Cross-modal hallucination
- This is a serious problem where multimodal AI systems generate details or information that are not actually supported or present when inputs combine different types of data, such as images and text.
- Adaptive Decoding
- MAD uses this concept to actively adjust the model's decoding process. Instead of generating tokens blindly, it dynamically selects the most reliable modality (e.g., visual or linguistic) for each step of the output generation.
- Multimodal Large Language Models
- These are AI systems capable of processing and understanding information from multiple types of input simultaneously, such as text, images, video, and audio.
Terminology
Summary
This paper presents Modality-Adaptive Decoding (MAD), a training-free method designed to mitigate cross-modal hallucinations
in Multimodal Large Language Models (MLLMs). These hallucinations occur when one modality inappropriately influences generation about another,
such as when visual cues cause a model to fabricate non-existent audio events. Addressing this is vital because it reveals a fundamental deficiency in modality-interaction control
that goes beyond simple single-modality errors.
The Limitation of Existing Methods
Current mitigation strategies, such as Audio-Visual Contrastive Decoding (AVCD), are often modality-agnostic,
meaning they lack awareness of task-specific requirements. These methods typically apply uniform distortions across all modalities without considering which input is actually relevant to the user's question. Consequently, a static approach cannot dynamically block irrelevant inputs from causing cross-modal corruption,
making it difficult to resolve complex interference patterns where one modality misleads the understanding of another.
How MAD Works
MAD introduces a task-driven modality weighting scheme
that adaptively scales contrastive decoding branches based on task requirements. Instead of using a fixed strength, MAD employs a four-branch formulation to aggregate distinct contrastive signals:
-
A joint audio-visual branch where both modalities are present.
-
A visual CD branch where audio is present but visual is perturbed.
-
An audio CD branch where visual is present but audio is perturbed.
-
Single-modality branches that fall back to contrastive decoding when one modality is absent to suppress interference from the remaining input.
By using these adaptive weights, the model can concentrate on relevant information while suppressing cross-modal interference.
Modality Self-Assessment
The core innovation of MAD is its ability to self-assess modality relevance
by querying the model itself. To determine the necessary weights, the system appends a fixed modality query prompt, such as: To answer this question, which modality is needed (audio, video, or both)?
The model's predicted logits for 'video', 'audio', and 'both' are then processed through a softmax function to extract normalized probabilities (w av, w v, and w a). This explicit modality awareness
allows the model to determine the importance of each sensory stream for every specific task without requiring additional training or supervision.
Experimental Results and Findings
Extensive testing on the CMM and AVHBench benchmarks demonstrates that MAD significantly reduces hallucinations, providing notable improvements for models like VideoLLaMA2-AV and Qwen2.5-Omni. The researchers found that MAD effectively addresses several types of dominance identified in the CMM benchmark:
-
Visual dominance, where the model over-relies on visual information at the expense of auditory or linguistic cues.
-
Audio dominance, where there is
excessive emphasis on auditory input.
-
Language dominance, occurring when models follow
linguistic priors even when they conflict with multimodal evidence.
This performance suggests that explicit modality-aware fusion
is a crucial component for robust multimodal reasoning.
Improvements for AI systems
Improvements
-
Inference-Time Modality-Adaptive Decoding (MAD) Integration: Implement a training-free decoding layer into the inference pipeline of Audio-Visual Large Language Models (AV-LLMs). This involves replacing standard greedy or uniform contrastive decoding with a four-branch weighted contrastive logit fusion mechanism.
-
Dynamic Self-Assessment Weighting Mechanism: Integrate a pre-generation
modality query
step where the model is prompted to identify the required modality (video, audio, or both) for a specific task. The system must extract these modality probabilities via softmax to derive task-specific weights (w v, w a, w av), which then dynamically scale the contrastive strength (alpha m) for each decoding branch. -
Multi-Branch Logit Fusion Architecture: Implement a specific logit aggregation formula that combines four distinct contrastive signals:
-
Joint Audio-Visual CD: To reinforce combined sensory grounding.
-
Visual CD (Audio Present): To suppress audio hallucinations triggered by visual cues.
-
Audio CD (Video Present): To suppress visual hallucinations triggered by auditory cues.
-
Single-Modality Fallback: To maintain performance when one modality is absent or perturbed.
Improved AI System Capabilities
-
Mitigation of Cross-Modal Hallucinations: The system will eliminate
video-driven audio hallucinations
(e.g., inventing the sound of splashing water just because a boat is visible) andaudio-driven visual hallucinations
(e.g., inventing a person dancing just because upbeat music is playing). -
Adaptive Modality Prioritization: The system will autonomously adjust its reasoning focus based on the query. For a question like
What color is the car?
, it will prioritize visual grounding; forWhat is the background music?
, it will prioritize auditory grounding, effectively suppressing interference from irrelevant modalities. -
Robustness Against Unimodal Bias: The system will resist
Visual Dominance
(over-reliance on sight at the expense of sound),Audio Dominance
(over-reliance on sound at the expense of sight), andLanguage Dominance
(relying on linguistic priors rather than actual sensory evidence). -
Zero-Shot Deployment: Because the improvement is implemented at the decoding level, the system can be upgraded to mitigate hallucinations across various pre-trained AV-LLMs (such as VideoLLaMA2-AV or Qwen2.5-Omni) without the need for expensive retraining or new annotated datasets.
Sources
- VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
- AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
- The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
- ImageBind-LLM: Multi-modality Instruction Tuning
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- MLVU: Benchmarking Multi-task Long Video Understanding
- Debiasing Multimodal Large Language Models via Penalization of Language Priors
- MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
- Do You Keep an Eye on What I Ask? Mitigating Multimodal Hallucination via Attention-Guided Ensemble Decoding
- Learning Transferable Visual Models From Natural Language Supervision
- Sigmoid Loss for Language Image Pre-Training
- HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
- Robust Speech Recognition via Large-Scale Weak Supervision
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection