From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs
cs.CL
Submitted: 2025-12-22
Updated: 2026-08-31
Comments: Paper accepted at ECML PKDD 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Subtitles are essential for video accessibility and audience engagement.
Terminology
Abstract
Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively reduced transcription errors on standard benchmark datasets. However, their performance in real-world production environments, particularly for non-English content like long-form Italian videos, remains largely unexplored. This paper presents a case study on developing a professional subtitling system for an Italian media company. To inform our system design, we evaluated four state-of-the-art ASR models (Whisper Large v2, AssemblyAI Universal, Parakeet TDT v3 0.6b, and WhisperX) on a 50-hour dataset of Italian television programs. The study highlights their strengths and limitations, benchmarking their performance against the work of professional human subtitlers. The findings indicate that, while current models cannot meet the media industry's accuracy needs for full autonomy, they can serve as highly effective tools for enhancing human productivity. We conclude that a human-in-the-loop (HITL) approach is crucial and present the production-grade, cloud-based infrastructure we designed to support this workflow.
Sources
- WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
- Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
- ASR Error Correction using Large Language Models
- End-to-End Speech Recognition: A Survey
- Robust Speech Recognition via Large-Scale Weak Supervision
- Anatomy of Industrial Scale Multilingual ASR
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- BLEURT: Learning Robust Metrics for Text Generation
- SubER: A Metric for Automatic Evaluation of Subtitle Quality
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering