UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca, Ethan Fetaya, Yftah Ziser, Gal Chechik, Haggai Maron
NVIDIA Research · Technion · Bar-Ilan University · University of Groningen
cs.CV, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Project Page: https://research.nvidia.com/labs/par/uniprobe/
Project page: https://research.nvidia.com/labs/par/uniprobe
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: UniProbe is a lightweight, unified, learnable detector for token-level hallucination detection in Large Vision-Language Models (LVLMs).
Terminology
Summary
UniProbe is a lightweight, unified, learnable detector for token-level hallucination detection in Large Vision-Language Models (LVLMs). It models a frozen LVLM's heterogeneous computational trace from a single forward pass, constructing a directed graph over image patches, query tokens, and generated tokens, with attention weights encoding their relations. UniProbe processes this trace with alternating structure-aware modules: a GNN for relational evidence, a ViT for 2-D visual geometry, and a GRU for response order. Interleaving these modules allows spatial, relational, and sequential evidence to interact throughout the detector.
The paper introduces a streaming variant for hallucination-aware decoding, which detects and resamples hallucinated tokens during generation, and a self-adaptation strategy aligning the detector with the LVLM's own generations. Across diverse LVLM backbones, UniProbe achieves state-of-the-art token-level and object-hallucination detection. During decoding, it reduces object hallucinations by up to 55% at 1.06× the latency of standard generation.
The method constructs a computational-trace graph from one forward pass, recording internal signals (layer-l hidden states and attention weights) and structural metadata (node type and patch coordinates for image nodes). Nodes are tokens of three types: response tokens, image patches, and query tokens. The graph is kept lightweight by scoring each image position by the total attention it receives from the response and retaining the highest-scoring positions; the same selection is applied to query tokens. Edges are derived directly from attention, retaining the strongest incoming connections for each response token from image patches, query tokens, and earlier response tokens.
The UniProbe architecture projects all node features into a shared representation space using type-specific linear maps, then stacks L identical blocks. Each block runs three modules in turn: a GNN that mixes evidence across modalities by updating every response node from its typed attention neighbors, a ViT over the image grid, and a BiGRU along the response in generation order. A linear head on the final response embeddings gives per-token hallucination probabilities.
For online detection, the bidirectional response GRU is replaced with a unidirectional one, ensuring that the score of each token depends only on the prefix. During decoding, if a token's hallucination probability exceeds a threshold τ, it is rejected and re-decoded. The paper also addresses distribution shift in self-generation by sampling a small subset of Objects365 images, generating free-form captions using the target model, and automatically labeling each generated object mention using CHAIR. The detector is then fine-tuned on these automatically labeled generations.
Experiments evaluate UniProbe on three setups: token-level hallucination detection (on MHALO and HalLoc benchmarks), object hallucination detection (on POPE), and hallucination-aware decoding during streaming generation (on COCO captions). On MHALO, UniProbe achieves F1M/F1IoU of 63.2/52.9 vs. 59.1/49.8 for HaloDet on GLM-4V, and 61.7/51.2 vs. 55.4/46.7 on Qwen-3-VL. On HalLoc, UniProbe beats the trained HalLocalizer on all three backbones. On POPE, UniProbe reaches 63.1 F1 / 90.0 AUC, compared to 41.0 / 75.0 for Token Grounding, the strongest prior detector. For streaming self-generated captions on COCO, the streaming detector achieves 32.6 F1obj before adaptation and 63.8 after self-adaptation, reducing CHAIRi/CHAIRs from 18.0/37.2 to 8.2/16.6, a 55% reduction, at only 1.06× vanilla latency.
Ablation studies show that flat MLP and Transformer probes perform substantially worse than the full model, trailing by at least 30.0 F1M. Removing any component hurts, with the largest drops caused by removing the relational graph (−10.1) and response recurrence (−4.6). The grounding signal lives in the middle layers, with detection peaking at layer 20 for GLM-4.1V and layer 14 for LLaVA-1.5, degrading toward the output. The decoding threshold τ trades hallucination reduction against content preservation, with τ = 0.70 being the most aggressive setting that reduces hallucinations without removing correctly generated content.
The paper concludes that structured internal model traces provide a practical signal for both hallucination detection and mitigation. Limitations include requiring access to hidden states and attention maps, applicability only to open-source models or directly by closed-source providers, and training a separate readout for each backbone as internal representations differ across models.
Improvements for AI systems
Improvements to AI systems:
-
Add a plug-and-play hallucination detector module to LVLMs – Integrate UniProbe as a lightweight post-hoc layer that reads the model’s own attention and hidden states during inference. The improved system can flag and correct token-level hallucinations in real time without retraining the base model, reducing object hallucinations by up to 55% at only 1.06× latency.
-
Implement hallucination-aware decoding with adaptive resampling – Replace greedy/beam decoding with a streaming loop that rejects any generated token whose UniProbe hallucination probability exceeds a threshold (e.g., τ=0.70) and re-decodes it. The improved system can generate captions or answers that are factually grounded in the image, preserving correct content while eliminating spurious objects.
-
Enable self-adaptation to distribution shift – Add an automatic fine-tuning pipeline that samples a small set of images, generates free-form captions with the target LVLM, labels object mentions via CHAIR, and fine-tunes UniProbe on those labels. The improved system can maintain high detection accuracy on the model’s own generation style, even when the base LVLM is fine-tuned or deployed on new domains.
-
Build a unified detector across multiple LVLM backbones – Use UniProbe’s architecture (GNN + ViT + GRU) as a shared readout that can be trained once per backbone but applied to any task requiring token-level grounding (e.g., visual QA, image captioning, referring expression comprehension). The improved system can serve as a universal safety layer for open-source LVLMs, flagging hallucinations before they reach the user.
-
Improve interpretability of LVLM outputs – Expose UniProbe’s per-token hallucination probabilities as a confidence score alongside each generated token. The improved system can provide users with a reliability metric, allowing them to trust or question specific claims, and enabling downstream systems to filter low-confidence tokens automatically.
-
Optimize computational trace usage for real-time applications – Leverage UniProbe’s lightweight graph construction (selecting top-attended image patches and query tokens) to reduce memory overhead. The improved system can run hallucination detection on edge devices or in streaming settings, making it feasible for real-time chatbots, assistive vision tools, or autonomous systems that require immediate factual grounding.
Abstract
Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input. Effective mitigation requires token-level localization, enabling targeted intervention without discarding the entire response. Existing detectors require expensive full-model fine-tuning, rely on external verifiers that ignore the model's generation process, or reduce internal signals to isolated features and hand-crafted statistics, discarding spatial, sequential, and relational structure. We introduce UniProbe, a lightweight, unified, learnable detector that models a frozen LVLM's heterogeneous computational trace from a single forward pass. UniProbe constructs a directed graph over image patches, query tokens, and generated tokens, with attention weights encoding their relations. It processes this trace with alternating structure-aware modules: a GNN for relational evidence, a ViT for 2-D visual geometry, and a GRU for response order. Interleaving them allows spatial, relational, and sequential evidence to interact throughout the detector. We further develop a streaming variant for hallucination-aware decoding, which detects and resamples hallucinated tokens during generation, and a self-adaptation strategy aligning the detector with the LVLM's own generations. Across diverse LVLM backbones, UniProbe achieves state-of-the-art token-level and object-hallucination detection. During decoding, it reduces object hallucinations by up to 55% at 1.06 times the latency of standard generation.
Sources
- Qwen3-VL Technical Report
- MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification
- Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
- VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
- Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models