Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters
cs.HC, cs.AI
Submitted: 2026-02-01
Updated: 2026-08-30
Comments: Accepted to ASSETS 2026
Code: https://github.com/Daria8976/MMAD
License: http://creativecommons.org/licenses/by/4.0/
The gist: Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded.
Terminology
Abstract
Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production, quality is rarely checked systematically. Existing evaluations rely on NLP metrics and short-clip guidelines, leaving open the question of how to assess long-form AD quality at scale. To address this, we developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth. Evaluations were based on a six-dimensional framework, grounded in professional guidelines and shaped by insights from our accessibility experts and blind consultants. Findings suggest that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters. However, qualitative analysis reveals that VLM reasoning is less reliable and actionable than that of human respondents. These insights underscore the potential of hybrid evaluation systems that leverage VLMs alongside human oversight, offering a path toward scalable AD quality control.
Sources
- SPICE: Semantic Propositional Image Caption Evaluation
- EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
- Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
- Qwen2.5-VL Technical Report
- Can Large Language Models Be an Alternative to Human Evaluations?
- LLM-AD: Large Language Model based Audio Description System
- CIDEr-R: Robust Consensus-based Image Description Evaluation
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
- The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
- Capturing Humans' Mental Models of AI: An Item Response Theory Approach
- Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
- VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
- Audio Description Customization
- GPT-4o System Card
- Movie Description
- Making Short-Form Videos Accessible with Hierarchical Video Summaries
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Towards Automatic Learning of Procedures from Web Instructional Videos
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support