Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-28
Comments: Under review
Project page: https://helsinki-nlp.github.io/shroom/2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (Shared-task on Hallucinations and Related Observable Overgeneration Mistakes in Vision language model s), which
Terminology
Abstract
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (Shared-task on Hallucinations and Related Observable Overgeneration Mistakes in Vision language model s), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA, image captioning, etc.). The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian. The shared task generated strong interest in the NLP community worldwide, with 27 teams contributing 600+ system submissions. The best systems achieve average scores of 0.58 in character-level correlation, 0.46 in label-conditioned correlation, and 0.51 in intersection-over-union (IoU) across four languages, outperforming the baselines by 30-40 points.
Sources
- Hallucination of Multimodal Large Language Models: A Survey
- Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- A Survey on Hallucination in Large Vision-Language Models
- DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Qwen3 Technical Report
- Visual Hallucination: Definition, Quantification, and Prescriptive Remediations
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering