Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking
Timothee Mickus, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Chuyuan Li, Aman Sinha, Lorenzo Vaiani, Jörg Tiedemann, Raúl Vázquez
cs.CV, cs.AI, cs.CL
Submitted: 2026-08-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated
Terminology
Abstract
In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.
Sources
- Hallucination of Multimodal Large Language Models: A Survey
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification
- Gemma 3 Technical Report
- A Survey on Hallucination in Large Vision-Language Models
- No Language Left Behind: Scaling Human-Centered Machine Translation
- H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models
- Qwen3 Technical Report
- Qwen3-VL Technical Report
- Visual Hallucination: Definition, Quantification, and Prescriptive Remediations
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models