Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs

arXiv:2609.17327 · cs.CL, cs.AI · Submitted 2026-09-15 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-09-15

Updated: 2026-09-16

Comments: Accepted to UncertaiNLP 2026 @ EMNLP. SHROOM-Visions 2026 shared task system description

Code: https://github.com/toqeerehsan/vlm_

Project page: https://helsinki-nlp.github.io/shroom/2026

License: http://creativecommons.org/licenses/by/4.0/

The gist: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages.

Terminology

Abstract

This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language models as independent annotators and combine their span predictions through character-level majority voting, and additionally explore activation probes. The approach ranks first in three of four languages and places on the podium in every language and metric. Our analysis indicates that disagreement among diverse models tracks disagreement among human annotators.

Sources

Related papers