Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs
cs.SD, cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/aarongrace/uncertainty2026
Terminology
Sources
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
- Language Models (Mostly) Know What They Know
- Visual hallucination detection in large vision-language models via evidential conflict
- VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation
- An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
- Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
- Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
- Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
- Hearing the Order: Investigating Position Bias in Large Audio-Language Models
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation
- ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
- All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen2-Audio Technical Report
- Qwen2.5-Omni Technical Report
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment