A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
Ali Shendabadi, Parnia Izadirad, Mostafa Salehi
Faculty of Intelligent Systems Engineering · College of Interdisciplinary Sciences and Technologies · University of Tehran
cs.CL, cs.AI, cs.LG, cs.SD
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: 6 pages
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 48/100
The gist: This research investigates "the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation." The study addresses the
Terminology
Summary
This research investigates the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation.
The study addresses the challenge that "the extracted frame-level representations are typically highdimensional—for example, Whisper-small produces 768-dimensional embeddings—leading to a large number of trainable parameters when followed by a trainable projection layer, which
often results in overfitting and inefficient training when emotion-labeled data are scarce."
To mitigate these issues, the authors propose a framework in which frame-level embeddings extracted from the Whisper encoder are reduced in dimensionality using PCA, eliminating the need for learned projection layers and substantially reducing the number of trainable parameters.
They replace the learned projection layer with PCA, an unsupervised dimensionality reduction technique that introduces no additional trainable parameters.
After reduction, the resulting embeddings are aggregated using an attention-based pooling mechanism
called QKV attention-based pooling mechanism,
which learns to assign higher weights to frames that are more informative for emotion recognition, enabling the model to focus on emotionally salient regions of an utterance.
The process concludes with a lightweight classification head
that maps the utterance-level embedding to emotion categories.
Additionally, the study investigates whether fine-tuning Whisper on a Persian automatic speech recognition (ASR) task improves downstream SER performance.
The researchers fine-tune Whisper-small on a Persian ASR dataset collected from YouTube... consisting of approximately 340 hours of speech,
which successfully reduces WER from 45% to 37%, indicating effective adaptation to Persian speech.
Experiments were conducted on the ShEMO dataset under a speaker-independent evaluation protocol.
The results indicate that PCA-based dimensionality reduction consistently improves emotion recognition performance while reducing training latency and memory usage.
Specifically, the PCA-based approach yields improvements on the order of 2% in Weighted Accuracy and 3% in Unweighted Accuracy when reducing the 768-dimensional Whisper-small representations to 512 dimensions.
Furthermore, PCA results in a substantial reduction in training latency—approximately 30% per epoch in our experiments—and lowers memory usage by eliminating additional trainable parameters.
Regarding language adaptation, the study finds that ASR fine-tuning yields only modest gains for SER, suggesting limited transfer from language adaptation to emotion-related representations under the evaluated conditions.
Specifically, ASR fine-tuning yields improvements of less than 1% in most cases.
The authors suggest this may be because SER systems trained solely on speech signals may rely more heavily on prosodic and paralinguistic cues, thereby reducing the impact of enhanced linguistic representations.
In conclusion, the proposed framework achieves state-of-the-art results on the ShEMO dataset, surpassing previously reported performance across both weighted and unweighted accuracy metrics.
Improvements for AI systems
1. Parameter-Efficient Feature Compression via Unsupervised Reduction
- What the improved system can do: By replacing heavy, trainable projection layers with PCA-based dimensionality reduction, the system can perform high-accuracy Speech Emotion Recognition (SER) on resource-constrained hardware (edge devices) and with extremely small datasets. It will prevent the overfitting typically seen when mapping high-dimensional embeddings (like Whisper’s 768-D) to emotion categories, while simultaneously reducing training latency by 30% and lowering memory overhead.
2. Salience-Aware Temporal Aggregation
- What the improved system can do: By implementing a QKV (Query-Key-Value) attention-based pooling mechanism, the system can move beyond simple frame averaging to
intelligent listening.
It will dynamically identify and assign higher weights to emotionally salient acoustic segments—such as sudden pitch shifts, pauses, or intensity changes—while ignoring emotionally neutral frames, leading to higher weighted and unweighted accuracy in detecting nuanced emotions.
3. Paralinguistic-Centric Language Adaptation
- What the improved system can do: Instead of relying solely on ASR (Automatic Speech Recognition) fine-tuning for language adaptation, the system will utilize a dual-stream fine-tuning approach that prioritizes prosodic and paralinguistic features. This ensures that when the model is adapted to a new language (e.g., Persian), it maintains or enhances its ability to detect how something is said (emotion) rather than just what is said (linguistic content), overcoming the diminishing returns of standard ASR-based transfer learning.
Abstract
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation. We propose a SER framework in which frame-level embeddings extracted from the Whisper encoder are reduced in dimensionality using PCA, eliminating the need for learned projection layers and substantially reducing the number of trainable parameters. The reduced representations are aggregated using an attention-based pooling mechanism and classified with a lightweight prediction head. In addition, we investigate whether fine-tuning Whisper on a Persian automatic speech recognition (ASR) task improves downstream SER performance. Experiments conducted on the ShEMO dataset under a speaker-independent evaluation protocol show that PCA-based dimensionality reduction consistently improves emotion recognition performance while reducing training latency and memory usage. ASR fine-tuning yields only modest gains for SER, suggesting limited transfer from language adaptation to emotion-related representations under the evaluated conditions. These findings provide practical insights into the efficient use of large pretrained speech models for emotion recognition in low-resource languages.
Sources
- Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
- EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering