Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach".
Jane: The paper was written by Yi-Cheng Lin, Huang-Cheng Chou and Hung-yi Lee from National Taiwan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everyone. Today we're digging into a paper with a real mouthful of a title: "Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach." Jane, I'm going to need you to break that down for me, because that's a lot of jargon in one sentence.
Jane: Happy to, Tom. So, the core problem is that when we build AI systems to recognize emotions from speech, they often work better for some groups of people than others. Think about it — a system might be great at detecting anger in male voices but miss it in female voices. That's what they call a "subgroup disparity."
Tom: And the "multi-label" part? That's throwing me off a bit.
Jane: That's actually a really important detail. Most older research treated emotion recognition as picking one emotion per sentence. But in reality, people often express multiple emotions at once. You can be both sad and angry, or happy and nervous. So this paper treats it as a multi-label problem, which is more realistic.
Tom: Okay, that makes sense. And the "pseudo-labeling and unsupervised learning" part? That's the fix they're proposing, right?
Jane: Exactly. The tricky part is that the usual way to fix bias is to know who's in which group — you need gender labels, age labels, race labels. But that information is often private or just not available. So they came up with a way to guess those groups without needing the actual labels.
Tom: Guessing? That sounds a little sketchy.
Jane: It's actually clever. They use two methods. One uses a pre-trained model to guess gender from the voice. The other just groups similar voices together using clustering, without any labels at all. Then they use those guessed groups to train the emotion system more fairly.
Tom: So it's like saying, "I don't know exactly who's in each group, but I have a good idea, and that's enough to make things fairer." I like that approach. It's practical.
Jane: And it works, too. They tested it on a standard dataset and got big improvements in fairness with only a small drop in overall accuracy. We'll get into the numbers in a bit.
Tom: I'm looking forward to that. Before we do, let me bring in Lu from Tsinghua. Lu, what's your first reaction to this idea of inferring demographics without asking for them?
Lu: I think it's a really smart workaround. The privacy angle is huge here. If you're building a voice assistant, you can't just ask users for their gender and race. That's invasive. But if you can infer enough about the group structure from the audio itself, you can apply fairness techniques without that burden. It opens up a lot of possibilities for real-world deployment.
Tom: So this isn't just an academic exercise. It's about making products that people actually use.
Lu: Absolutely. And the fact that they're using a multi-label setup makes it even more relevant, because that's how emotions actually work in the real world.
Tom: Alright, I'm hooked. Let's get into the details of how they actually built this thing.
Summary: Jane: So, Tom, let's talk about the actual method. The paper introduces something called the Implicit Demography Inference module, or IDI. It's a fancy name for a simple idea: figure out the groups without asking for labels.
Tom: And they do that two ways, right? One with a pre-trained gender detector, and one with clustering.
Jane: Exactly. The first way is pseudo-labeling. They take a pre-trained model that's already good at detecting gender from speech, and they use it to label all their training data. That gives them a guess for each speaker's gender.
Tom: And the second way is unsupervised clustering. They take a different pre-trained model — one that's good at identifying speakers — and they extract a kind of "voiceprint" from each utterance. Then they run a clustering algorithm called K-Means on those voiceprints to group similar voices together.
Lu: And the nice thing about that second approach is that it doesn't assume anything about what the groups mean. It just finds natural groupings in the data. Those groupings might correspond to gender, but they might also pick up on age or accent or other vocal characteristics.
Tom: So it's a more general approach. It's not just about gender.
Lu: Right. And that's why they tested it on race and age too, not just gender. The clustering method doesn't care what the demographic attribute is — it just finds the structure.
Jane: Once they have these inferred groups, they use them to train the emotion recognition model with fairness constraints. They tried four different debiasing techniques: reweighting, downsampling, and two versions of something called Distributionally Robust Optimization.
Tom: Those all sound like ways to make sure the model doesn't just get good at recognizing emotions for the majority group.
Jane: Exactly. Reweighting gives more importance to samples from underrepresented groups. Downsampling just removes some samples from the majority group to balance things out. And the robust optimization methods try to minimize the worst-case loss across all groups.
Tom: So they're not just hoping the model figures it out — they're actively pushing it to be fair.
Jane: Right. And the results were pretty impressive. With pseudo-labeling, they improved fairness metrics by over twenty-eight percent while only dropping accuracy by less than two percent. With the unsupervised clustering, they got over twenty percent improvement in one fairness metric with about a four point seven percent drop in accuracy.
Meng: I want to jump in here. That accuracy drop — is that acceptable in practice? For a real product, a four point seven percent drop in emotion recognition accuracy is noticeable.
Jane: That's a fair question, Meng. The paper argues it's a trade-off. You're sacrificing a little bit of overall performance to get much fairer behavior across groups. And in many applications, that's worth it.
Meng: I guess it depends on the application. For a customer service bot, maybe. For a medical diagnosis tool, that drop might be too much.
Lu: But it's worth noting that the fairness improvements are substantial. A twenty percent improvement in fairness is a big deal, especially when you're getting it without any demographic labels.
Tom: And that's the real story here. You're getting meaningful fairness gains without needing the sensitive data that usually makes this kind of work impossible in practice.
Jane: Exactly. And we're going to dig into what that means for real-world applications in a minute.
Improvements: Tom: So we've talked about what the paper does, but what does it actually improve over what came before? Jane, what's the state of the art before this?
Jane: Good question. Most prior work on debiasing speech emotion recognition relied on having explicit gender labels. You'd train a model, see where it's biased, and then use those labels to correct it. But that doesn't work when you don't have the labels.
Tom: And there were some unsupervised methods from computer vision, right? The paper mentions a few.
Jane: Right, they compared against methods called LfF and DisEnt. Those are designed for image classification, and they don't transfer well to speech. The paper shows that those methods barely improve fairness at all — they actually just hurt performance without fixing the bias.
Lu: That's a common problem in this field. Methods that work for images don't always work for audio, because the data structure is so different. Speech has temporal dynamics, speaker characteristics, all kinds of things that images don't have.
Tom: So this paper is saying, "Hey, we need methods designed for speech, not just borrowed from vision."
Jane: Exactly. And that's what they're offering. Their unsupervised clustering method, in particular, is designed for speech because it uses a speaker verification model to extract embeddings that capture vocal characteristics.
Meng: I'm curious about the practical side. How much compute does this require? Running a pre-trained model to extract embeddings, then clustering, then training a separate emotion model — that sounds like a lot of steps.
Jane: They mention using two Nvidia V100 GPUs and about five hundred GPU hours total. So it's not trivial, but it's also not crazy for a research project.
Meng: Okay, that's manageable. And the fact that they're using pre-trained models means you don't have to train everything from scratch. That's a big win.
Lu: And there's another improvement worth highlighting. The paper evaluates on the full test set without filtering out samples that don't have a clear emotion label. A lot of prior work would just remove those ambiguous samples, which makes the problem easier. This paper keeps everything, which gives a more honest picture of performance.
Tom: So they're not cherry-picking the easy cases. They're dealing with the messy reality of real speech.
Jane: Right. And that makes the results more trustworthy. If you're getting good fairness numbers on the full test set, that's more meaningful than getting great numbers on a cleaned-up subset.
Tom: And the fact that the unsupervised method also improves fairness for race and age, not just gender — that's a big deal. It means the approach generalizes.
Lu: It does. And that's what makes me excited about the potential here. If you can infer group structure without labels, you can apply this to any demographic attribute, or even to attributes you haven't thought of.
Meng: So the real improvement here is a method that's practical, doesn't need sensitive data, and works across multiple types of bias. That's a solid contribution.
Tom: I agree. And I want to hear what Lalam thinks about where this could go next.
Conclusion: Tom: So, Lalam, you've been listening to all of us geek out over this paper. What's your take on where this research leads?
Lalam: I think the most exciting implication is cultural. Speech emotion recognition is being used in education, in healthcare, in customer service. If these systems are biased, they're going to treat people unfairly in those settings. A student with a non-native accent might have their frustration misread as anger. A patient with a speech disorder might have their anxiety missed entirely.
Jane: That's a really important point. The bias isn't just an academic problem — it has real consequences for how people are treated.
Lalam: Exactly. And this paper offers a way to reduce that bias without requiring the kind of demographic data that would make people uncomfortable. That means it could actually be deployed in the real world, not just in a lab.
Tom: So we're talking about making voice assistants and therapy tools and educational software more fair for everyone, without asking users to reveal sensitive information about themselves.
Lalam: That's the vision. And the fact that the unsupervised method works across gender, race, and age suggests it could be a general tool for fairness in speech systems.
Meng: I want to add a practical note. The paper is honest about its limitations. It's tested on one dataset, CREMA-D, which is acted emotions, not natural speech. So we need to see if this works in the wild.
Jane: That's a fair caveat. But it's a strong first step. And the fact that they're using the full test set, not a filtered version, gives me more confidence in the results.
Tom: Alright, let's wrap this up. The paper we've been discussing, "Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach," shows that you can reduce bias in emotion recognition without needing demographic labels. You can guess the groups from the audio itself, and that's enough to make the system fairer.
Lu: And it's a meaningful improvement over prior work, both in terms of fairness gains and in terms of not needing sensitive data.
Tom: Exactly. It's a practical solution to a real problem, and it opens the door for more research in this direction.
Jane: And with that, we're going to say goodbye to this paper. Thanks for joining us, everyone. We'll be back next time with another exciting piece of research.
Tom: Until then, keep listening, keep learning, and keep asking questions. See you on the next episode.
Yi-Cheng Lin, Huang-Cheng Chou, Hung-yi Lee
National Taiwan University
eess.AS, cs.CL, cs.SD
Submitted: 2025-05-30
Updated: 2026-08-18
Comments: Accepted by InterSpeech 2025. 7 pages including 2 pages of appendix
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 83/100
Key concepts
- Subgroup Disparity
- This refers to situations where an AI system for speech emotion recognition performs better for some groups of people than others. For example, a system might accurately detect anger in male voices but fail to do so in female voices.
- Multi-label Problem
- Instead of recognizing only one emotion per sentence, this approach treats the task as multi-label, acknowledging that people often express multiple emotions simultaneously, such as being both sad and angry.
- Implicit Demography Inference (IDI)
- This module infers demographic groups without asking for labels. It uses two methods: a pre-trained model to guess gender from voice and unsupervised clustering of voiceprints to group similar speakers together.
- Fairness Constraints
- The inferred groups are used to train the emotion recognition model with fairness constraints. Techniques like reweighting and distributionally robust optimization actively push the model to be fair across different demographic groups.
Terminology
Summary
Summary
This paper addresses the underexplored issue of fairness and subgroup disparities in categorical Speech Emotion Recognition (SER), specifically in the multi-label setting. The authors note that existing debiasing methods in SER predominantly rely on explicit demographic annotations (e.g., gender, age, race), which are difficult to obtain due to privacy concerns and are scarce in real-world data. They also highlight that prior approaches are designed for single-label categorical SER, whereas recent research emphasizes the importance of multi-label SER, aligning with psychological studies on emotion co-occurrence.
To overcome these limitations, the authors introduce an Implicit Demography Inference (IDI) module that simulates demographic supervision by inferring latent group cues directly from speech data, without requiring manual demographic labels. The IDI module operates through two techniques:
-
Pseudo-labeling: Leveraging a pre-trained gender detection model (audeering/wav2vec2-large-robust-24-ft-age-gender) to generate proxy demographic labels for each input utterance. The model's reliability is assessed on the CREMA-D dataset, achieving an accuracy of 94.4%.
-
Unsupervised clustering: Extracting embeddings using the ECAPA-TDNN model and applying K-Means clustering (with cluster sizes of 2, 4, 8, 16, and 32) to uncover inherent group structures. The cluster assignments are treated as group labels, reflecting latent structures that may correspond to demographic differences.
The inferred group labels are then incorporated into debias training of the SER model, using four debiasing methods: Reweighting (RW), Downsampling (DS), Group Distribution Robust Optimization (GDRO), and Group Aware DRO (GADRO). These methods impose fairness constraints to decrease performance disparities across the inferred subgroups.
The experiments are conducted on the CREMA-D database, using a WavLM base plus feature extractor with two linear layers as the SER backbone. The authors simulate biased training data by setting a gender imbalance ratio of 1:20 (female to male or male to female), while evaluating models on the original test set to reflect real-world conditions. They use macro-F1 score and Hamming accuracy for SER performance, and Equal Opportunity (TPRgap) and Demographic Parity (DPgap) for fairness metrics.
Key results:
-
Pseudo-labeling IDI reduces subgroup disparities, improving fairness metrics (TPRgap and DPgap) by over 28% (specifically 28.78% and 29.13%, respectively) with less than a 2% decrease in SER accuracy (a 4.38% drop in F1 and 1.37% drop in ACC).
-
Unsupervised IDI (with K=16 clusters) yields more than a 4.6% improvement in fairness metrics (20.05% in TPRgap and 4.61% in DPgap) with a drop of less than 3.6% in SER performance (a 4.76% decrease in F1 and 3.55% decrease in ACC).
-
The unsupervised IDI method consistently mitigates race and age disparities, demonstrating its generalizability when explicit demographic information is unavailable. For race and age, the unsupervised method with K=16 clusters improves the TPRgap metric by over 12.75% while maintaining an SER performance decrease of 4.76% and 3.55% in F1 and ACC, respectively.
-
Compared to existing unsupervised debiasing methods (LfF and DisEnt), the proposed approach exhibits a more favorable trade-off between fairness and SER performance. While earlier methods yield minor performance improvements, they often incur few debias effects in TPRgap and DPgap.
-
The number of clusters plays a critical role; a cluster size of 16 provides balanced granularity, offering significant fairness improvements while maintaining minor performance drops.
The authors also visualize ECAPA-TDNN embeddings using t-SNE, confirming that male and female samples are clearly separated, validating that the clustering captures gender information.
Limitations acknowledged by the authors: The experiments are based on the CREMA-D dataset (acted emotional expressions), so generalizability to naturalistic settings remains to be validated. The methods depend heavily on the quality of pre-trained models (gender detection model for pseudo-labeling and ECAPA-TDNN for embedding extraction).
Conclusion: The paper demonstrates that pseudo-label learning improves TPR and DP metrics by 28.78% and 29.13%, respectively, with only a minor 4.38% degradation in SER accuracy. The proposed unsupervised method achieves 20.05% and 4.61% improvements in TPR and DP metrics, respectively, with just a 4.76% drop in SER performance. Future work will explore more possible bias attributes like content, language, or speaking style.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:
Improvement: Integrate the Implicit Demography Inference (IDI) module into any SER system. This module uses:
-
Pseudo-labeling from a pre-trained gender/age detector (e.g.,
wav2vec2-large-robust-24-ft-age-gender) to generate proxy demographic labels. -
Unsupervised clustering (K-Means, k=16) on ECAPA-TDNN embeddings to discover latent group structures.
Then, apply any of the four debiasing techniques (Reweighting, Downsampling, GDRO, GADRO) using these inferred labels instead of ground-truth demographics.
Resulting capability: The SER system reduces gender-based subgroup disparities by 28.78% (TPR gap) and 29.13% (DP gap) with pseudo-labels, and by 20.05% (TPR gap) and 4.61% (DP gap) with unsupervised clustering, while losing less than 4.8% in F1 score. This works even when no demographic annotations are available.
Improvement: Adopt the paper’s training strategy that simulates real-world biased data (e.g., 1:20 gender imbalance) while evaluating on the original, unbiased test set. This prevents overoptimistic performance estimates common in prior debiasing work.
Improvement: Replace single-label SER with multi-label classification using a threshold of 1/Y (e.g., 1/6 for six emotions). Combine this with the debiasing losses (RW, DS, GDRO, GADRO) that operate per emotion class and per inferred subgroup.
Improvement: Use the paper’s finding that K=16 clusters is optimal for ECAPA-TDNN embeddings. Additionally, set λ GD=4 for GADRO to balance worst-case loss and group-size regularization.
Improvement: Adopt the paper’s evaluation metrics: TPR gap (Equal Opportunity) and DP gap (Demographic Parity), computed as RMS across all emotion classes. Use macro-F1 and Hamming accuracy for performance.
Improvement: Leverage the finding that unsupervised clustering (without any demographic labels) reduces bias not only for gender but also for race and age. This is achieved by training the debiasing module on cluster assignments that correlate with multiple demographic axes.
Improvement: Use the paper’s backbone (WavLM base + two linear layers) and the IDI module with pre-trained models (ECAPA-TDNN, gender detector). The total training time is 500 GPU hours on two V100s, making it feasible for most research labs.
-
Recognize multiple emotions per utterance (multi-label) with fairness guarantees across gender, race, and age.
-
Operate without any demographic annotations, using only speech features.
-
Maintain high accuracy (F1 ≈ 0.64, ACC ≈ 0.82) while reducing subgroup bias by up to 41% (TPR gap) compared to a standard ERM model.
-
Generalize to unseen demographic groups and biased training distributions, making it suitable for real-world, privacy-sensitive applications like virtual assistants, mental health monitoring, and customer service analytics.
Abstract
While subgroup disparities and performance bias are increasingly studied in computational research, fairness in categorical Speech Emotion Recognition (SER) remains underexplored. Existing methods often rely on explicit demographic labels, which are difficult to obtain due to privacy concerns. To address this limitation, we introduce an Implicit Demography Inference (IDI) module that leverages pseudo-labeling from a pre-trained model and unsupervised learning using k-means clustering to mitigate bias in SER. Our experiments show that pseudo-labeling IDI reduces subgroup disparities, improving fairness metrics by over 28% with less than a 2% decrease in SER accuracy. Also, the unsupervised IDI yields more than a 4.6% improvement in fairness metrics with a drop of less than 3.6% in SER performance. Further analyses reveal that the unsupervised IDI consistently mitigates race and age disparities, demonstrating its potential when explicit demographic information is unavailable.
Sources
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions