OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
cs.CL
Submitted: 2026-03-25
Updated: 2026-08-31
Terminology
Sources
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Baichuan-Omni Technical Report
- OmniBench: Towards The Future of Universal Omni-Language Models
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
- VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
- GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
- EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
- Learning Transferable Visual Models From Natural Language Supervision
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering