Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
cs.CL, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-16
Comments: Accepted to UncertaiNLP 2026 @ EMNLP. SHROOM-Visions 2026 shared task system description
Code: https://github.com/toqeerehsan/vlm_
Project page: https://helsinki-nlp.github.io/shroom/2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages.
Terminology
Abstract
This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language models as independent annotators and combine their span predictions through character-level majority voting, and additionally explore activation probes. The approach ranks first in three of four languages and places on the podium in every language and metric. Our analysis indicates that disagreement among diverse models tracks disagreement among human annotators.
Sources
- Hallucination of Multimodal Large Language Models: A Survey
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- Gemma 4 Technical Report
- Qwen3-VL Technical Report
- Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
- A Survey on Hallucination in Large Vision-Language Models
- Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking
- Overthinking Causes Hallucination: Tracing Confounder Propagation in Vision Language Models
- VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
- Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
- Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering