Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding
cs.SD, cs.AI, eess.AS
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/zhenghuatan/rVAD
Terminology
Sources
- The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems
- WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
- Qwen2-Audio Technical Report
- Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
- TAC: Timestamped Audio Captioning
- AffectGPT-R1: Leveraging Reinforcement Learning for Open-Vocabulary Multimodal Emotion Recognition
- Understanding R1-Zero-Like Training: A Critical Perspective
- The Devil is in the EOS: Sequence Training for Detailed Image Captioning
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
- Speech Emotion Diarization: Which Emotion Appears When?
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- SECap: Speech Emotion Captioning with Large Language Model
- Qwen3 Technical Report
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment