How to Reduce Whisper Hallucination
cs.CL
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/malaysia-ai/dataset
Terminology
Sources
- Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
- FMA: A Dataset For Music Analysis
- FSD50K: An Open Dataset of Human-Labeled Sound Events
- The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
- Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- LoRA: Low-Rank Adaptation of Large Language Models
- Careless Whisper: Speech-to-Text Hallucination Harms
- Granary: Speech Recognition and Translation Dataset in 25 European Languages
- Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
- Earnings-22: A Practical Benchmark for Accents in the Wild
- SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering