Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
cs.CL, cs.AI, cs.SD
Submitted: 2026-04-20
Updated: 2026-09-13
Comments: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics
Code: https://github.com/SALT-NLP/CAVA
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly.
Terminology
Abstract
The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy. Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benchmark scores. To understand how well these scores align with what practitioners ultimately care about, user satisfaction, we collect 776 human preference ratings from realistic voice assistant conversations, finding that both subsets and full benchmark achieve only 0.85 correlation with human. To better predict preferences, we trained regression models on these selected subsets, achieving 0.98 correlation -- outperforming regression models trained on both random subsets and the full benchmark. This demonstrates that in regression modeling, well-curated subsets outpredict the full benchmark, showing quality over quantity. We open-source these regression-weighted subsets as the HUMANS benchmark, an efficient proxy for LAM evaluation that captures both benchmark performance and user preferences.
Sources
- Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Mind the Gap! Static and Interactive Evaluations of Large Audio Models
- PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
- Voxtral
- AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
- tinyBenchmarks: evaluating LLMs with fewer examples
- AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
- How Benchmark Prediction from Fewer Data Misses the Mark
- WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering