Dual-Pathway Circuits of Object Hallucination in Vision-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Dual-Pathway Circuits of Object Hallucination in Vision-Language Models".
Jane: Vision-language models often produce unreliable outputs through object hallucinations, and this study proposes a framework to mechanistically understand these errors by identifying distinct computational circuits within diverse VLMs.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, this paper, "Dual-Pathway Circuits of Object Hallucination in Vision-Language Models," proposes a framework called Dual-Pathway Circuit Analysis to identify circuits specifically related to hallucinations in VLMs. The central thesis is that there's a shared dual-pathway organization within these models that drives the errors we see when they hallucinate objects.
Jane: What the authors claim is pretty specific, Tom; they apply activation patching across five different, architecturally diverse VLMs to find two distinct pathways: one that supports correct predictions and another pathway that drives those erroneous outputs.
Lu: It’s really interesting that they found this dual-pathway organization is consistent across all five models tested, suggesting it's a general structural property of how these models process information rather than something specific to just one architecture.
Meng: If the structure is shared, that means we might be able to find universal ways to address these hallucination errors without needing a custom fix for every single model we deploy.
Lalam: That shared structure could give us a much better blueprint for improving the AI culture by making sure our models adhere to more consistent visual grounding standards across different applications.
Tom: Exactly, and they use Conditional Pathway Analysis, or CPA, to dig deeper into how these two pathways interact at a component level. The main finding there is that while both pathways show redundancy within each pathway subset, the grounding pathway exhibits a specific polarity flip between correct and hallucinating samples.
Jane: That polarity flip is the key thing they highlight; on correct samples, most of the grounding components have positive effects, but when we look at hallucinating samples, that flips to negative.
Conclusion: Tom: So, looking at the paper "Dual-Pathway Circuits of Object Hallucination in Vision-Language Models," it seems the authors are arguing that object hallucination stems from a specific, consistent dual-pathway organization inside these models, which they've managed to map out using activation patching and CPA.
Jane: What this means in simple terms is that the model has two functional routes for processing visual information: one that leads to correct answers based on the image, and another route that consistently pushes it toward generating things that aren't actually present.
Lu: The implication here is pretty big for how we think about these systems; if we can isolate this hallucination pathway, we might be able to target and suppress the errors without severely hurting the model's ability to perform its primary visual tasks.
Meng: From a practical standpoint, if they can scale down or modify just that one erroneous component group, it suggests a way to make our deployed systems significantly more reliable for tasks where accuracy matters most.
Lalam: I think this points toward developing much more reliable vision tools where we can be confident in what the AI is actually seeing versus what it's imagining.
Tom: That’s the big picture, Jane; they’ve shown that object hallucination isn't just some random glitch but a predictable circuit behavior within the model's architecture. This research provides a clear mechanistic roadmap for understanding and potentially controlling this failure mode in vision-language models.
UIUC (University of Illinois Urbana-Champaign) · UMich (University of Michigan) · Stanford University · HKUST (Hong Kong University of Science and Technology) · PKU (Peking University) · NUS (National University of Singapore)
cs.CV
Submitted: 2026-05-13
Updated: 2026-10-07
Importance score: 85/100
The gist: Vision-language models often produce unreliable outputs through object hallucinations, and this study proposes a framework to mechanistically understand these errors by identifying distinct
Key concepts
- Dual-Pathway Organization
- This refers to the finding that all tested vision-language models share two distinct computational routes. One route is responsible for correctly identifying objects (grounding), while the other route is responsible for generating incorrect or hallucinated outputs. These pathways exist regardless of the model's specific architecture.
- Visual Grounding Pathway
- This pathway consists of components that are crucial for accurately connecting visual information to correct predictions. Components in this pathway show a positive effect when the model is correct, indicating their role in accurate object recognition and understanding.
- Hallucination Pathway
- This pathway drives erroneous outputs, leading to object hallucinations. Components here show a negative effect during hallucination, meaning they are more active or influential when the model produces incorrect answers.
Terminology
Summary
Vision-language models often produce unreliable outputs through object hallucinations, and this study proposes a framework to mechanistically understand these errors by identifying distinct computational circuits within diverse VLMs. The core finding is that object hallucination in VLMs arises from a shared dual-pathway organization where grounding components exhibit a consistent polarity flip between correct and hallucinating samples, while the hallucination pathway lacks such a flip.
Dual-Pathway Circuit Discovery
The researchers applied activation patching across five architecturally diverse VLMs to identify a visual grounding pathway that supports correct predictions and a hallucination pathway that drives erroneous outputs.
This initial analysis revealed a consistent dual-pathway organization across all five models.
The visual grounding pathway is characterized by components with an indirect effect (IE) larger when the model answers correctly, indicated by a Cohen’s d > 0. Conversely, the hallucination pathway consists of components with a negative Cohen’s d (< 0), meaning their indirect effect is larger during hallucination.
Pathway-Level Interaction Analysis
To characterize how these pathways interact, Conditional Pathway Analysis (CPA) was introduced to move beyond single-component analysis. CPA identifies whether components are redundant or synergistic within each pathway by comparing the joint effect of restoring a pathway component against the sum of its individual effects. The diagnostic metric MagDiff(P)
is used to quantify this: negative values indicate magnitude-redundancy,
while positive values indicate magnitude-synergy.
The analysis revealed that both pathways are redundant within each subset, as the magnitude ratio IE(P)/IE(Ci) stays in the range of [0.07, 0.69].
Grounding Pathway Polarity Flip
A critical finding from CPA is that the grounding pathway additionally exhibits a consistent IE polarity flip between subsets.
On correct samples, 61–83% of grounding-pathway component IEs are positive,
but on hallucinating samples, this flips: only 21–33% of IEs remain positive, and the subset-mean IE turns negative (−1.08 to −0.07) in every model.
This shift is a signature invisible to standard component-level analysis because it involves the interaction between components across subsets.
Causal Validation and Intervention
The study validated the circuit hypothesis through targeted suppression of hallucination-pathway components, showing that scaling these components reduces object hallucination by up to 76% with minimal accuracy cost.
The intervention strategies compared include:
-
Uniform scaling of the halluciation pathway outputs (Equation 5).
-
Top-k component selection, which is necessary for models with
dispersed circuits,
identifying acompact core of strongly hallucination-driving components.
-
Targeted interventions like Inferencetime Intervention (ITI) and mean-difference projection baselines.
Cross-Architecture Generalization
The analysis confirmed that the dual-pathway organization is not architecture-specific but reflects a shared mechanistic structure that develops during training rather than being inherited from a specific architecture.
At the macro level, the spatial distribution of components is consistent across models, with hallucination components concentrating in early layers and at network boundaries
and grounding components in mid-to-late layers.
Furthermore, the identified circuits transfer selectively: they are effective against object-existence questions but show type-dependent transfer
to relational hallucination on the AMBER benchmark. The asymmetry observed across different hallucination types suggests that attribute hallucination involves partially distinct circuits not captured by the POPE-derived pathway.
Conclusion
The overall conclusion is that the hallucination pathway acts through multiple distinct activation directions rather than a single shared one,
as evidenced by the geometric analysis of the singular spectrum. The study concludes that grounding ablation destroys task performance, while hallucination-pathway suppression reduces hallucination at near-zero accuracy cost, confirming that the two pathways carry functionally distinct information. The results support a shared dual-pathway organization
for object-existence hallucination across five architecturally diverse VLMs.
The gist
Activation patching reveals a consistent dual-pathway organization across all studied models, consisting of a visual grounding pathway and a hallucination pathway which may lead to errors.
Table 11: Models and experimental setup.
Model Arch. type LLM backbone Layers Sig. Grnd. Hall. rate
Qwen3-VL-8B Embed-concat Qwen3 (8B) 36/26/14/12/13.1%
LLaVA-v1.6-7B Projector-concat Mistral (7B) 32/48/18/30/13.4%
Llama-3.
Improvements for AI systems
Here are the specific improvements that can be made to existing Vision-Language Models (VLMs) based on the findings of this paper, and what those improved systems can achieve:
The core improvement lies in moving from black-box mitigation strategies (like training alignment or post-hoc correction) to a mechanistic, circuit-level intervention. The proposed framework allows for targeted suppression of specific neural pathways responsible for hallucination.
Here are the specific improvements and resulting capabilities:
-
The ability to perform
Pathway Suppression
on identified hallucination circuits via component-wise scaling during inference (Section 5.5). -
The use of Top-K Component Selection (Section 5.5) instead of uniform scaling when the hallucination circuit is highly dispersed, allowing for targeted pruning of the most critical error-driving components.
-
The development of a robust, architecture-agnostic diagnostic tool: Conditional Pathway Analysis (CPA), which reveals that grounding components exhibit a consistent polarity flip between correct and hallucinated samples—a signature missed by single-component analysis.
-
The use of Logit Lens Analysis to validate the functional roles assigned to these circuits, confirming whether grounding layers are indeed image-faithful and hallucination layers are pushing toward language-prior errors.
AI Systems with These Improvements Can Do:
-
A model can be modified at inference time by selectively scaling (suppressing) specific activation components identified as part of the
hallucination pathway.
This allows the system to reduce object hallucination by up to 76% while maintaining minimal accuracy cost (as shown in Table 13). -
If a model's error-driving circuit is too complex or dispersed, the system can dynamically select only the top-K components with the largest negative causal effect during inference. This provides a more precise and efficient mitigation strategy than uniform scaling alone.
-
The system gains a deeper understanding of why errors occur. By analyzing the CPA signature, researchers can confirm that while grounding components are redundant (carrying overlapping information), they undergo a critical polarity flip on hallucinated samples, suggesting that the model is not simply missing visual evidence but has been entrained to output its non-visual language-prior belief when it hallucinates.
-
The system's internal architecture can be functionally validated against external benchmarks. Logit Lens Analysis confirms that specific layers are responsible for grounding (image fidelity) versus hallucination (language bias), ensuring that the interventions are targeting the correct functional components, leading to more reliable and interpretable mitigation techniques across diverse VLM architectures.
In summary, these improvements transform VLMs from opaque systems into mechanistically understood
systems capable of self-correction or targeted error reduction based on their internal computation structure.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- Video-R1: Reinforcing Video Reasoning in MLLMs
- Localizing Model Behavior with Path Patching
- Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
- Aligning Large Multimodal Models with Factually Augmented RLHF
- Steering Language Models With Activation Engineering
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- Mitigating Image Captioning Hallucinations in Vision-Language Models
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models