AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?
cs.SD, cs.AI
Submitted: 2026-06-19
Updated: 2026-08-26
Terminology
Sources
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language Models
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
- Zero-shot audio captioning with audio-language model guidance and audio context keywords
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- DeepSeek-V3 Technical Report
- Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
- Step-Audio 2 Technical Report
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment