Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization

arXiv:2608.29924 · cs.CV, cs.AI, cs.LG · Submitted 2026-08-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization".

Jane: The paper was written by Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu and Prashant Khanduri from Department of Computer Science, Wayne State University and Institute for AI and Data Science (AIDaS), Wayne State University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary & Core Mechanism: Tom: We’ve established that these models tend to hallucinate, and we know this paper aims to fix that by stabilizing the internal features. Jane, can you explain the core mechanics of this approach in simple terms?

Jane: Think of it like a rigorous internal quality control system for the AI's thoughts. When an image comes in, the model generates many hidden representations—its "features." The authors’ method ensures that these features stay tightly coupled to that original visual input throughout the whole process, preventing them from drifting into incorrect conclusions.

Tom: So, they aren't just checking the final output word; they are stabilizing the internal thought process itself? That’s a huge conceptual leap.

Jane: Precisely! It's about making sure that every step of every feature is consistent with the visual input, ensuring that if something is there in the picture, its representation inside its strongly reflects it.

Meng: This sounds like a form of highly specialized regularization designed for visual grounding, which is far more targeted than general training techniques we usually apply to large models. It’s focused on physical reality within the image space.

Lu: I agree with Meng; it's about stabilizing the *meaning* derived from vision at every step, not just adding an external check on what it says at the end. That provides a much deeper level of foundational strength for us to build upon.

Lalam: From a societal view, this internal stability means that when we use AI for interpreting complex visual data—say, reading an old map or analyzing a historical document—the resulting narrative will be much closer to factual evidence.

Tom: It sounds like they’ve built a deep safety net for the model’s internal thinking process, but how does this compare to existing methods that intervene at inference time? Lu, what does this internal approach imply for how we view AI reliability moving forward?

Lu: It implies that future models can move beyond just statistical correlation; they are being trained to achieve a kind of visual causality. They aren't just guessing what usually goes with what; they are confirming *why* it must be there based on the evidence.

Jane: And the beauty of the summary is that this process happens implicitly, meaning it’s woven into the feature space rather than bolted on as a massive, external module we have to add at inference time.

Meng: That sounds much more scalable for deployment. If we can stabilize features inherently, we minimize latency and keep the model's size manageable while boosting its factual accuracy during training.

Lalam: This reliability means that if AI-generated educational materials are used, historical narratives will carry a far greater weight of truth, making learning less prone to subtle misinformation.

Tom: It sounds like they’ve built a much deeper safety net for the model’s internal thinking process. Now that we know how it works, let's look at the specific results and improvements in Segment three.

Improvements & Advanced Details: Tom: So, we’ve seen how INFUSE works by building stability directly into the model's internal features instead of adding external checks at inference time. Jane, what are the most impressive quantitative outcomes of this approach?

Jane: The authors report that across LLaVA and Qwen models, this implicit stabilization cuts down on hallucination significantly, which is a huge relief for anyone using these large models.

Meng: The quantitative data is what really strikes me—they're talking about reductions of forty-six to sixty-three percent on the AMBER dataset for LLaVA-one point five. That level of reduction in factual errors is astonishingly high, Meng thinks.

Lu: I think it’s worth noting how consistent this performance is across all three different backbones. That suggests the method isn' robust and doesn't just work optimally for one specific architecture or model type, which is a major win for scalability.

Lalam: It also looks like the improvements aren’t at the expense of general performance, which is great news for culture. The model maintains its ability to perform visual question answering accurately while becoming far more reliable in a way that benefits society.

Tom: Exactly, Jane; it doesn't sacrifice accuracy in VQA-v2 or TextVQA. That’s a huge win for real-world deployment where reliability is paramount.

Meng: And the computational cost is very low—only about six point two GPU hours on just two thousand samples. That’s incredibly practical when compared to other methods that require massive amounts of data or complex pipelines to achieve similar results in practice.

Lu: The theoretical justification, especially the anchor's contraction rate at /K, suggests a deep mathematical guarantee of consistency we’re seeing in practice. It's not just luck; there' a formal reason for this stability.

Jane: That rigor explains why we aren't seeing unexpected instability when dealing with perturbed inputs anymore, as the system is designed to be much more robust to noise and ambiguity than before.

Lalam: This level of reliability means that if AI is used for complex tasks like reading medical scans or interpreting technical manuals, the potential for misinformation drops drastically.

Tom: It seems like a genuinely powerful way to build trust in the AI output. Now, let's look at how this translates to real-world scenarios and specific benchmarks in Segment four.

Paper discussion segment 3: Tom: We’ve seen how INFUSE works by building stability directly into the model's internal features. The paper also details its robustness under visual corruption. Jane, can you explain what the team found when they introduced perturbations like masking or noise?

Jane: They tested the model with various visual distortions, such as removing parts of an image or adding slight noise to it, and found that while small distortions cause predictable drops in accuracy—which is expected—the model’ hallucinations remain impressively low compared to its baseline.

Tom: That makes sense; you can't expect a perfect response if half the picture is missing. But the interesting finding seems to be how well it handles different types of corruption, doesn's it?

Meng: The results show that corruptions which preserve the overall structure, like slight rotation or cropping, are actually quite well tolerated by this trained model. It’s not just robust to noise; it’s robust to maintaining global context while changing local details.

Lu: This is where the theoretical work really pays off because of the stabilization—the internal feature space has become so concentrated that small perturbations can't push the output distribution far enough to change the probability of a hallucination.

Lalam: For cultural applications, this means we don't have to discard images just because they are slightly blurry or poorly framed; AI can still derive accurate information from imperfect visual data, which is a huge boon for accessibility.

Tom: It’s not just robust to noise; it’s also shown that the gains in object grounding—like F1A and F1R scores—are actually improved compared to existing methods. That's a very subtle but important detail.

Meng: And critically, this method maintains high performance across benchmarks like MMHal-Bench, which is designed to test complex reasoning, proving that we aren’re not sacrificing deep understanding for better stability.

Lu: The implications are that we are seeing a shift from merely "correct" answers to something genuinely robust and persistent in the the high-dimensional space of visual information.

Jane: This level of reliability means that if AI is used for complex tasks like reading technical manuals, the potential for misinformation drops drastically.

Tom: It seems like a genuinely powerful way to build trust in the AI output, knowing it can handle imperfect inputs while maintaining its accuracy. Now, let's look at how this fits into the larger landscape of AI in Segment five.

Conclusion: Tom: We’ve seen that by weaving stability directly into the model’s core features through INFUSE, we’re making a massive leap forward for AI reliability. It sounds like we've covered all major aspects of the paper today.

Jane: It’s truly a huge step toward building a more trustworthy vision-language system, especially given how much hallucination has plagued these powerful models in the past.

Meng: I just hope this becomes the standard way to train these models because it offers such an efficient path to high performance while remaining computationally practical for real-world deployment at scale.

Lu: I think the theoretical grounding—the mathematical proof of anchor contraction—is what we need, ensuring consistent results not just a set of quick fixes or tweaks.

Lalam: This stability means that even when AI assists us with complex tasks like reading historical documents or interpreting visual art, it will be more faithful to the evidence we provide.

Tom: I agree; it moves us toward a future where the AI acts as a dependable tool rather than just a sophisticated guess.

Jane: The fact its operational cost is low also makes this an extremely practical solution for various applications, which is something that deserves serious consideration in practice.

Meng: And it doesn' not require any extra overhead during deployment, which is crucial for real-world implementation and avoids adding latency to the entire system.

Lu: It feels like we’re seeing a shift from merely "correct" to something genuinely robust and persistent in the high-dimensional space of information.

Lalam: This ensures that the cultural exchange facilitated by AI is grounded in reality, preventing subtle, pervasive misinformation across all mediums.

Tom: A huge relief for everyone involved in creating or consuming this technology. We've covered everything from the core mechanism to its real-world impact, so we're ready to move on to a new topic for the next segment of our show. This is a fascinating piece of work by Sarker and their team titled "Hallucination Mitigation for Large Vision-Language Models via Implicit Feature Stabilization."

Aditi Sarker, Rafi Ibn Sultan, Hui Zhu, Dongxiao Zhu, Prashant Khanduri

Department of Computer Science, Wayne State University · Institute for AI and Data Science (AIDaS), Wayne State University

cs.CV, cs.AI, cs.LG

Submitted: 2026-08-30

Updated: 2026-08-30

Comments: 28 Pages, 12 Figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: This paper introduces INFUSE (Implicit Feature Stabilization and Alignment), a training framework designed to mitigate hallucinations in Large Vision-Language Models (LVLMs).

Key concepts

Implicit Feature Stabilization
This core mechanism ensures that a model's internal representations (features) remain tightly coupled to the original visual input throughout processing. It acts as an internal quality control system, making sure every step of the feature is consistent with the image evidence.
Hallucination Mitigation
The technique significantly reduces factual errors in AI models by preventing them from drifting into incorrect conclusions. This stability allows for more reliable interpretation of complex visual data, such as historical documents or medical scans.
Visual Grounding
This is a highly targeted approach focused on physical reality within an image space. It ensures the model's derived meaning is based directly on the evidence present in the picture, rather than relying on general statistical correlations.

Terminology

Summary

This paper introduces INFUSE (Implicit Feature Stabilization and Alignment), a training framework designed to mitigate hallucinations in Large Vision-Language Models (LVLMs). The authors address the issue of feature instability, where mild semantics-preserving perturbations of the input cause large changes in the learned embeddings, a phenomenon that correlates with increased hallucination rates. By building perturbation-invariance... into the model weights during fine-tuning, INFUSE offers a way to improve reliability without the inference-time overhead required by explicit intervention methods.

The Problem of Feature Instability

The researchers identify a connection between hallucination and a measurable property of their representations, feature instability. In existing LVLMs, semantically equivalent observations of a scene often produce substantially different visual features when subjected to controlled perturbations like additive noise, masking, or partial region removal. The authors distinguish their approach from existing explicit remedies, which impose stability from outside the trained model at inference time through methods like latent-space steering or constrained decoding. Unlike these methods, which pay a cost on every query, INFUSE utilizes implicit stabilization where perturbation-invariance is built into the model weights during fine-tuning.

How INFUSE works

INFUSE is a two-stage training framework that first stabilizes representations within each modality before aligning them across modalities. The ordering is deliberate: alignment acts on representations that are already low-variance, so positive pairs are matched on semantics rather than on appearance noise.

The framework operates through the following stages:

  • Stage 1 ([S1] Intra-modal stabilization): This stage stabilizes visual and textual representations. Visual embeddings are pulled toward perturbation-averaged anchors and repelled from [a] lagged reference embedding to prevent degenerate optima. Textual embeddings are pulled toward human-corrected truthful-response anchors and repelled from hallucinated-caption embeddings across the batch.

  • Stage 2 ([S2] Inter-modal alignment): This stage aligns the stabilized representations across modalities with bidirectional contrastive objectives (V to T and T to V) plus a standard generation loss.

Theoretical and Empirical Results

The framework is supported by a formal justification involving a variance-contraction guarantee and a Lipschitz argument connecting embedding stability to stability of the output. Specifically, the anchor’s root-mean-square deviation from the perturbation-mean representation shrinks at rate 1/sqrt K in the number of views.

Empirical evaluations on LLaVA-1.5, LLaVA-1.6, and Qwen3-VL-8B-Instruct demonstrate significant improvements:

  • INFUSE reduces AMBER CHAIR by 46-63% relative to each base model.

  • It improves performance on ObjHal, MMHal, HallusionBench, and POPE.

  • It preserves VQA-v2 and TextVQA, ensuring no degradation in general vision-language understanding.

  • The method is cheap to train and requires nothing extra running at deployment.

Improvements for AI systems

Improvements

  1. Implementation of a Two-Stage Stabilize-then-Align Training Framework (INFUSE):
  • Stage 1: Intra-modal Stabilization:

  • Visual Branch: Construct a low-variance visual anchor by generating K=100 masked views of an input image (using a 97% patch-mask ratio) and averaging their embeddings. Train the model to pull current visual embeddings toward this perturbation-averaged anchor while applying a lagged-parameter repulsion term (repelling the current embedding from the representation produced by the previous optimization step) to prevent representation collapse.

  • Textual Branch: Train the language model to pull the embeddings of its own generated responses toward human-corrected, truthful response anchors and simultaneously repel them from the embeddings of rejected/hallucinated captions.

  • Stage 2: Inter-modal Alignment:

  • Apply bidirectional InfoNCE (CLIP-style) objectives to align the stabilized visual and textual representations. This ensures the model matches semantic meaning rather than aligning to transient appearance noise or unstable features.

  1. Parameter-Efficient Weight-Level Integration:
  • Integrate these stabilization objectives using Low-Rank Adaptation (LoRA) with a rank of 16 and a scaling factor of 32, specifically targeting the multimodal projector and the final four transformer layers of the Large Language Model (LLM).

  • Keep the vision encoder frozen during both stages to prevent the re-introduction of encoder-level instability during cross-modal alignment.

Improved AI System Capabilities

  • Zero-Latency Hallucination Mitigation: The system will drastically reduce object-existence hallucinations (e.g., describing non-existent objects) and attribute/relation errors (e.g., incorrect colors or spatial positions) without requiring any inference-time interventions, latent steering, or constrained decoding, thereby maintaining peak deployment speed.

  • High Robustness to Visual Corruption: The system will maintain consistent semantic embeddings and reliable outputs even when presented with perturbed inputs, including heavy Gaussian noise, significant partial occlusions (masking), or geometric transformations (rotation and cropping).

  • Superior Fine-Grained Multimodal Reasoning: The system will demonstrate significantly higher accuracy in complex reasoning tasks, including precise object counting, geometric/spatial reasoning, and the accurate interpretation of structured visual data such as charts, tables, and maps.

  • Preservation of General Intelligence: The system will mitigate hallucinations while maintaining or improving baseline performance on general vision-language understanding benchmarks like VQA-v2 and TextVQA.

Abstract

Large Vision-Language Models (LVLMs) are prone to hallucinations: they fluently describe objects, attributes, and scenes that are not in the image. We connect part of this failure to a measurable property of their representations, feature instability, where mild semantics-preserving perturbations of the input cause large changes in the learned embeddings; hallucination rates rise together with this variability. Existing stability-motivated remedies are explicit, in the sense that they intervene at inference time through latent steering or constrained decoding, and pay for it on every query. We propose implicit stabilization instead: perturbation-invariance is built into the model weights during fine-tuning, and nothing extra runs at deployment. Our framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives. We prove that the anchor's root-mean-square deviation from the perturbation-mean representation shrinks at rate 1/sqrt K in the number of views, and that under a Lipschitz decoder, this bounds how much any perturbation can change the model's hallucination behavior. On LLaVA-1.5, LLaVA-1.6, and Qwen3-VL-8B-Instruct, INFUSE reduces AMBER CHAIR by 46-63% relative to each base model, improves ObjHal, MMHal, HallusionBench, and POPE, and preserves VQA-v2 and TextVQA, all with no inference-time overhead.

Sources

Related papers