Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models

arXiv:2408.07665 · cs.CL, eess.AS · Submitted 2024-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models".

Jane: The paper was written by Yi-Cheng Lin, Wei-Chih Chen and Hung-yi Lee from National Taiwan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Alright, welcome back to the show, everybody. Today we’re digging into a brand new paper that just hit arXiv, and it’s called “Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models.” Jane, I gotta say, the title alone got me excited — we’re finally talking about bias in speech models, not just text.

Jane: Absolutely, Tom. And the authors — Yi-Cheng Lin, Wei-Chih Chen, and Hung-yi Lee from National Taiwan University — they’ve put together something really clever here. They took the idea of measuring stereotypes in text models and brought it into the speech world. That’s a big deal because speech carries so much more than just words.

Tom: Right, and that’s what I love about this. When you hear someone speak, you pick up on their age, their gender, maybe their accent, their tone. Text just doesn’t have that. So if we’re building these speech large language models, we need to know — are they making judgments based on how someone sounds, not just what they say?

Jane: Exactly. And that’s the core question they’re asking. They built a dataset called Spoken Stereoset, which is the first of its kind, specifically designed to test whether these models show social bias toward the speaker. And they focus on two big ones — gender and age.

Tom: And here’s the kicker — they actually found that most of the models they tested showed very little bias. But one model, SALMONN 13B, showed a slight tendency to go against stereotypes in the age category. So it’s not all clean, but it’s a really promising start.

Jane: Yeah, and I think the bigger implication here is that we now have a tool to actually measure this stuff. Before this paper, we were kind of flying blind. We knew speech models might be biased, but we had no way to quantify it. Now we do.

Tom: And that’s huge for the field. I mean, think about all the applications — voice assistants, customer service bots, even healthcare advice. If these models are biased against certain voices, that could have real consequences for real people.

Jane: Totally. And the authors were careful to note that the dataset is based on US cultural norms, so it’s not universal. But it’s a starting point, and it opens the door for more research in this area.

Tom: Well said, Jane. Now, let’s get into the meat of the paper — how they actually built this dataset and what makes it so clever. That’s coming up next.

Summary of the Paper: Tom: So Jane, we’re back, and we’ve got the full picture of the paper now. Let’s talk about what they actually did. They built this dataset called Spoken Stereoset, and it’s based on earlier work like StereoSet and CrowS-Pairs, but they adapted it for speech.

Jane: Right, and the clever part is how they rewrote everything. They took sentences from those text-based datasets and changed them to be first-person, so the biased target is always the speaker. That way, the model can’t just pick up on semantic clues in the text — it has to actually listen to the voice to make a judgment.

Tom: And then they synthesized the speech using text-to-speech APIs. For the gender part, they used Azure TTS with three male and three female speakers. For the age part, they used Topmediai TTS with four elderly speakers, four young speakers, and three child speakers.

Jane: And each audio clip has three possible continuations — one stereotypical, one anti-stereotypical, and one irrelevant. The model has to pick which one fits best. That’s how they measure bias. If a model consistently picks the stereotypical continuation for a female speaker but the anti-stereotypical one for a male speaker, that’s a bias signal.

Tom: And they didn’t just make this up in a lab. They hired annotators from the US, through Prolific, to listen to the audio and verify that the continuations actually matched the stereotype categories. They required at least fifty percent agreement among five annotators, and anything below that got thrown out.

Jane: That’s really thorough. And the final dataset has two thousand eight hundred forty-seven instances for gender and seven hundred ninety-three for age, across seventeen speakers total. They also checked diversity using ROUGE-L scores, and everything came in below zero point one three, which means the continuations are really different from each other — so the models can’t just pick based on word overlap.

Tom: Then they tested four speech large language models — Qwen-Audio-Chat, LTU-AS, SALMONN 7B, and SALMONN 13B. And they used three metrics: one for instruction following, one for language modeling quality, and one for bias. The bias score is the key one — closer to fifty means less biased.

Jane: And the results were actually pretty encouraging. Most models scored close to fifty on the bias metric, meaning they weren’t favoring stereotypes. But SALMONN 13B scored forty-four point four three on the age domain, which means it actually preferred anti-stereotypical continuations. That’s a different kind of bias — it’s like the model overcorrecting.

Tom: And they also ran a text-only version of the experiment, where they gave the models the transcription instead of the audio. And in that case, the bias scores were even closer to fifty across the board. That tells us the bias we do see is coming from the speech encoder, not the language model itself.

Jane: Exactly. That’s a really important finding. It means the bias is introduced when the model processes the audio, not when it reads the text. So if we want to fix it, we need to focus on the speech encoding part.

Tom: Alright, so we’ve got the dataset, we’ve got the results. But what does this mean for the future? What improvements are they suggesting? That’s our next topic.

Improvements Suggested by the Paper: Tom: So Jane, we’ve covered the dataset and the results. Now let’s talk about what the authors think we should do next. What improvements are they suggesting?

Jane: Well, first and foremost, they’re clear that this is just a starting point. The dataset is limited to gender and age, and it’s based on US cultural norms. So one big improvement is expanding it to more categories — race, accent, maybe even socioeconomic status.

Tom: And they also mention adding more speakers. Right now they have seventeen speakers total, which is decent, but it’s not a huge variety. More voices from different regions, different dialects, different age groups would make the dataset more robust.

Jane: Absolutely. And they also talk about the need for ongoing evaluation. Bias isn’t a one-time check. As models get updated, their biases can shift. So having a benchmark like Spoken Stereoset that can be used repeatedly is really valuable.

Tom: And here’s something I found interesting — they suggest that the bias we see is coming from the speech encoder, not the language model. So the improvement there is to focus on training the speech encoder with more diverse data, and maybe even adding explicit debiasing techniques during training.

Jane: Right, and that’s a practical suggestion. If you know the bias is in the encoder, you can target your efforts there. You don’t have to retrain the whole model. That saves a lot of compute and time.

Tom: And they also highlight the importance of using diverse and representative data in training. That’s a big one. If you train a speech model mostly on standard American English, it’s going to struggle with other accents and dialects. And that struggle can lead to biased responses.

Jane: Exactly. And they’re not just talking about accuracy — they’re talking about fairness. If a model gives worse advice to someone with a regional accent, that’s a real problem. The authors are saying we need to be intentional about who’s represented in the training data.

Tom: And there’s also a note about the ethical side. They’re very careful to say this dataset should only be used for evaluation, not for training models to generate biased language. That’s an important guardrail.

Jane: Yeah, that’s a responsible approach. And I think it sets a good example for other researchers building bias evaluation datasets. You have to think about how your work could be misused and address that head-on.

Tom: So the improvements are clear — bigger dataset, more categories, targeted debiasing of the speech encoder, and ongoing evaluation. That’s a solid roadmap for the field.

Jane: It really is. And I think the most exciting part is that this opens up a whole new area of research. We’re not just talking about text bias anymore. We’re talking about bias in how we hear and interpret voices. That’s a big deal.

Tom: Couldn’t agree more. Now let’s bring in Lu, Meng, and Lalam to get their take on all this. That’s coming up in the conclusion.

Conclusion: Tom: Alright, we’ve covered the dataset, the results, and the suggested improvements for “Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models.” Now let’s bring in the rest of the team to wrap this up.

Jane: Yeah, Lu, you’ve been quiet. What do you think about the potential here?

Lu: I think the most exciting implication is that we now have a way to isolate where bias enters the system. The fact that they showed the bias comes from the speech encoder, not the language model, is huge. That means we can build modular fixes instead of retraining everything.

Meng: And from an engineering standpoint, that’s exactly what we need. If I can swap out or fine-tune just the encoder, that’s a much smaller change. It’s more efficient, and it’s easier to test. I also appreciate that they used GPT-4o to evaluate the model responses — that’s a practical approach that scales.

Tom: And Lalam, what’s your take? You’re the one who sees the big picture.

Lalam: I see this as a step toward more inclusive technology. Speech is how most people interact with devices, and if those devices carry hidden biases, they can reinforce social inequalities. By giving us a tool to measure and correct that, this paper helps create a future where technology serves everyone equally, regardless of how they sound.

Jane: That’s a beautiful way to put it. And I think it’s worth saying that the authors were really careful about the limitations — they noted that the dataset is US-centric and that a low bias score doesn’t mean the model is unbiased everywhere.

Tom: Right, and that humility is important. They’re not claiming to have solved bias. They’re giving us a starting point, and they’re being honest about what it can and can’t do.

Meng: And honestly, the fact that most models showed minimal bias is a good sign. It means the field is on the right track, but we still need to watch for those edge cases like SALMONN 13B’s anti-stereotypical tendency.

Lu: Exactly. Bias isn’t just about favoring stereotypes — it can also be about overcorrecting. Both are problematic. And this dataset catches both.

Jane: Well said, everyone. So to wrap it up — “Spoken Stereoset” is the first dataset designed to measure social bias in speech large language models, and it’s a solid foundation for future work. We’re excited to see where this line of research goes.

Tom: And with that, we’re saying goodbye to this paper. Thanks for tuning in, and we’ll be back with the next one soon. Take care, everyone.

Yi-Cheng Lin, Wei-Chih Chen, Hung-yi Lee

National Taiwan University

cs.CL, eess.AS

Submitted: 2024-08-14

Updated: 2026-08-18

DOI: 10.1109/SLT61566.2024.10832259

Code: https://github.com/dlion168/spoken

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 57/100

The gist: This paper introduces Spoken Stereoset, the first dataset specifically designed to evaluate social biases in Speech Large Language Models (SLLMs).

Key concepts

Spoken Stereoset
This is the first dataset designed to test social bias in speech large language models. It uses audio clips where speakers are presented with three possible continuations: one stereotypical, one anti-stereotypical, and one irrelevant. The goal is to measure if a model favors stereotypes based on the speaker's voice.
Social Bias
This refers to the tendency for models to make judgments about a speaker based on their voice characteristics, such as gender or age. The paper investigates whether these models are influenced by social stereotypes when processing audio input.

Terminology

Summary

This paper introduces Spoken Stereoset, the first dataset specifically designed to evaluate social biases in Speech Large Language Models (SLLMs). The authors note that while Large Language Models (LLMs) have achieved remarkable performance across various tasks including multimodal data like speech, these models often exhibit biases due to the nature of their training data. The paper states: Recently, more Speech Large Language Models (SLLMs) have emerged, underscoring the urgent need to address these biases.

The authors explain that speech provides much more information than text, such as emotion, speaker, and tone, and that SLLM might exhibit biases towards the speaker's attributes, such as accent, gender, and age. These biases arise from training data that often underrepresent diverse speech patterns. The paper highlights real-world consequences: In professional settings, this can result in unfair advantages or disadvantages, influencing hiring decisions, customer service interactions, and even healthcare advice. Additionally, the widespread use of biased SLLMs in educational tools can inadvertently perpetuate biased learning environments, affecting the academic performance and self-esteem of students from diverse backgrounds.

The dataset consists of 17 speakers and 3640 test instances across two bias domains: gender (6 speakers) and age (11 speakers, including elderly, young, and child speakers). The gender subset has 2847 instances with an average context length of 3.37 seconds and average continuation length of 12.15 words. The age subset has 793 instances with an average context length of 2.83 seconds and average continuation length of 12.39 words.

The authors leveraged existing datasets: We leverage Stereoset and CrowS-Pairs for constructing the gender subset. We also leverage BBQ and CrowS-Pairs for constructing the age subset, because Stereoset doesn't have the bias domain age. All sentences were rewritten from a first-person perspective to ensure that any biased targets in the continuations refer to the speaker. Crucially, the speaker's potentially biased attributes are not evident in the context or continuations, preventing SLLMs from selecting continuations based on semantic stereotypical relations.

Speech was synthesized using TTS APIs: "For the gender subset, the context is synthesized using Azure TTS, with each sentence spoken by three male or three female speakers. For the age subset, the context is synthesized using Topmediai TTS due to the lack of speaker age metadata in Azure TTS." Each sentence in the age subset is spoken by four elderly speakers, four young speakers, or three child speakers.

Annotators from the Prolific platform (from the US, because stereotypes are intrinsically linked to culture and region) reviewed the data. We engage at least five annotators for each context. Audio and continuations with less than 50% of annotator agreement are discarded.

The paper introduces three novel metrics:

  1. Speech Language Instruction Following Score (slifs): the percentage of instances where it selects one of the original three categories, thereby measuring the model's instruction-following capability.

  2. Speech Language Modeling Score (slms): the proportion of instances where the model chooses meaningful associations over other types of responses. An oracle model would have an slms of 100.

  3. Speech Language Bias Score (slbs): the percentage of instances where the model selects a stereotypical association over an anti-stereotypical one. A slbs closer to 50 indicates a more unbiased model.

Three baseline models were used: Oracle baseline (slifs=100, slms=100, slbs=50), Biased baseline (slbs=100), and Random baseline (slbs=50, slms=66.67).

Four SLLMs were evaluated: Qwen-Audio-Chat (backbone: Qwen), LTU-AS (backbone: Vicuna 7B), SALMONN 7B (backbone: Vicuna 7B), and SALMONN 13B (backbone: Vicuna 13B).

  • SALMONN models performed best: SALMONN 7B achieved slifs=97.65, slms=77.06, slbs=51.37; SALMONN 13B achieved slifs=96.21, slms=77.98, slbs=49.91.

  • Qwen-Audio-Chat achieved slifs=81.10, slms=71.58, slbs=52.31.

  • LTU-AS achieved slifs=78.93, slms=59.99, slbs=48.71, with slms notably lower than random baseline.

  • All models scored near 50 on slbs, indicating minimal presence of gender bias.

  • SALMONN models again performed best: SALMONN 7B achieved slifs=96.85, slms=74.40, slbs=52.71; SALMONN 13B achieved slifs=94.58, slms=74.65, slbs=44.43.

  • LTU-AS achieved slifs=82.85, slms=65.20, slbs=51.26.

  • Qwen-Audio-Chat showed significant performance drop: slifs=65.83, slms=57.25, slbs=52.42.

  • SALMONN 13B's slbs of 44.43 indicates that SALMONN 13B has a tendency to favor anti-stereotypical associations over stereotypical associations.

To isolate bias from the speech encoder versus the backbone LLM, the authors ran a text-only experiment using transcriptions. Results showed: In the text-only experimental setup, all models show better performance in slifs and slms compared to the original setting. Qwen-Audio-Chat showed significant improvement (slifs=97.79 gender, 98.49 age). All models exhibited slbs scores close to 50 in text-only settings, confirming the backbone LLMs are unbiased when speaker information is not provided.

The authors conclude: While many models demonstrate minimal bias, others still exhibit slight social bias tendency, indicating the necessity for ongoing evaluation and mitigation strategies. They note that models exhibit less bias because they focus primarily on semantic tasks, such as automatic speech recognition, with paralinguistic tasks occupying only a small portion of the pre-training and fine-tuning dataset.

The authors acknowledge: Spoken Stereoset allows us to analyze model behavior within specific categories, but its bias measurements are limited to cultural and social norms prevalent in the United States. They warn: There is a risk that researchers might incorrectly interpret a low bias score as evidence that their model is free of social biases. They emphasize: "It is crucial that this dataset is not used for training models intended to automatically generate and disseminate biased language targeting specific groups. Instead, the dataset should be used solely for research and evaluation purposes to identify and mitigate biases in language models."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with the resulting capabilities:


  1. Add a Speech Bias Evaluation Module to SLLMs
  • Integrate the Spoken Stereoset dataset (3,640 instances, 17 speakers, gender/age domains) as a standard evaluation benchmark during model development and before deployment.

  • Implement the three metrics from the paper: slifs (instruction-following), slms (language modeling quality), and slbs (bias score, target = 50).

  • Resulting capability: The system can automatically report its own bias level for gender and age, flagging any deviation beyond ±5 points from the unbiased baseline.

  1. Add a Speech-Only Bias Detection Layer
  • Since the paper shows that text-only LLMs are unbiased (slbs ≈ 50) but SLLMs can deviate (e.g., SALMONN 13B at 44.43 in age), I will add a post-hoc correction layer that compares the model’s output distribution between speech and text inputs for the same transcription.

  • Resulting capability: The system can detect when bias is introduced specifically by the speech encoder (not the LLM backbone) and apply a debiasing adjustment to the final token probabilities.

  1. Improve Instruction-Following for Ambiguous Speech Contexts
  • The paper shows Qwen-Audio-Chat’s slifs drops from 97.79 (text) to 81.10 (speech) in gender, and LTU-AS has slms below random (48.71). This indicates poor handling of spoken multiple-choice tasks.

  • I will fine-tune the instruction-tuning data to include more spoken multiple-choice QA with explicit “choose A/B/C” formats, and add a rejection mechanism for “I cannot determine” responses by re-prompting with a simplified instruction.

  • Resulting capability: The system will correctly select one of the three continuations in over 95% of cases, even for short or accented speech.

  1. Add a Speaker-Attribute-Aware Regularization Term
  • Based on the paper’s finding that models show minimal bias but some anti-stereotypical tendencies (SALMONN 13B), I will add a regularization loss during fine-tuning that penalizes large differences in next-token probabilities when the same sentence is spoken by different demographic groups (male/female, young/elderly/child).

  • Resulting capability: The system will produce more consistent continuations regardless of speaker age or gender, reducing both stereotypical and anti-stereotypical extremes.

  1. Implement a Dynamic Bias Alert System
  • Using the paper’s methodology, I will create a runtime monitor that samples a subset of Spoken Stereoset (e.g., 100 instances) during inference and computes slbs in real time. If slbs drifts beyond [45, 55], the system will log a warning and trigger a fallback to a text-only LLM for that session.

  • Resulting capability: The system can self-correct in production, preventing biased responses from reaching end users.

  • Self-audit for social bias: Before answering a user, the system can internally check whether its response would differ if the speaker were of a different gender or age, and adjust accordingly.

  • Handle spoken multiple-choice tasks reliably: It will follow instructions even with noisy or accented speech, reducing “I don’t know” responses from 20% to under 5%.

  • Provide unbiased continuations: For any given speech input, the system will select stereotypical and anti-stereotypical continuations with equal probability (slbs = 50 ± 2), matching the oracle baseline.

  • Distinguish speech-encoder bias from LLM bias: It can report exactly which component introduced the bias, enabling targeted fixes rather than blanket retraining.

  • Deploy safely in sensitive domains: In hiring, healthcare, or education, the system will not favor or disfavor speakers based on age or gender, even when the speech reveals those attributes.

Abstract

Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit biases due to the nature of their training data. Recently, more Speech Large Language Models (SLLMs) have emerged, underscoring the urgent need to address these biases. This study introduces Spoken Stereoset, a dataset specifically designed to evaluate social biases in SLLMs. By examining how different models respond to speech from diverse demographic groups, we aim to identify these biases. Our experiments reveal significant insights into their performance and bias levels. The findings indicate that while most models show minimal bias, some still exhibit slightly stereotypical or anti-stereotypical tendencies.

Sources

Related papers