Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations

arXiv:2505.18466 · cs.CL · Submitted 2026-05-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations".

Jane: The paper was written by Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, Ziwei Zhu and George Mason University from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary & Intersectional Analysis: Jane: The researchers found that the bias isn’t just one thing; it’s a combination of identity markers. They use this sophisticated concept of intersectionality to see how things like religion and marital status pile on top of each other to create specific, often harmful, stereotypes.

Tom: It's not just one factor causing the bias; it's the way they found that combining gender with childbearing expectations creates a unique kind of bias that is completely missed by standard evaluation metrics.

Lu: This is where their specialized lexicon comes into play—that’s a huge leap forward because you can’t measure what you don't have named. They created this culturally grounded lexicon, which has nine hundred twenty-three terms, to capture these specific cultural stigmas. It’s truly brilliant how they identified those concepts that were previously unmeasured.

Meng: That lexicon allows us to quantify the bias using a method called Bias TF-IDF. It essentially lets us assign a numerical weight to how much these specific biased terms appear in the generated text for different identities, which is incredibly useful for practical auditing.

Lalam: This helps us see that the LLM isn't just making random statements; it is actively linking certain groups, like single Muslim women, with terms related to social seclusion. This reveals how our current AI might be inadvertently reinforcing negative cultural expectations in its output.

Tom: And this pattern of bias persists across all three applications they tested: storytelling, hobbies, and to-do lists. It’s not limited to one type of text generation; it’s woven throughout the model's output, which is a major discovery.

Lu: I think the biggest revelation here is how subtle these biases are, especially when we look at everyday tasks like creating a to-do list for a character. They're not always screaming; they're just whispering harmful stereotypes that are hard to catch.

Meng: The data shows that these intersectional harms are quantifiable and measurable using this TF-IDF score across the forty-eight different identity combinations they tested. It gives us an actual number to work with, which is incredibly practical for development teams looking at real-world impact.

Lalam: It allows us to see the weight of cultural expectations in a way that lets us fix the vision, knowing exactly where we need more nuance and less stereotype in our AI's representation of South Asia.

Jane: So, we’ve seen how they measure it—through this sophisticated intersectional analysis; now let’ explore their strategies to try and correct the output in our next segment.

Improvements & Mitigation: Tom: The core of this research is looking at two ways to make the LLMs less biased—simple and complex debiasing prompts. We want to see how effective these tools are, particularly across those different language families.

Jane: The paper presents these as attempts to correct the original, biased output, but it’s important not to think of them as a cure; they're a tool for mitigation. They aren't perfect, and the findings are very instructive about where they fall short when we look at cultural nuance.

Lu: I find the comparison between simple and complex debiasing really interesting because of how much the results vary by language family. It’s clearly not a one-size-fits-all solution, which is a major insight for global AI deployment.

Meng: The practical takeaway here is that simple prompts have limited power, especially in Indo-Aryan texts where the biases are deeply ingrained. We need those more specific, complex instructions to even see any meaningful change in the output.

Lalam: This suggests that if we want our models to serve diverse cultures fairly, we can't just throw a generic instruction at them; we have to be culturally specific in how we prompt and train them to achieve balance.

Tom: And it’s not just about English-centric fixes; the researchers found that while complex prompts help in some Dravidian contexts, they simply don't work as effectively in Indo-Aryan languages. That was a surprising finding for many of our listeners.

Lu: This points to a real technical challenge—the cultural bias is so deeply encoded in those specific linguistic datasets that standard debiasing techniques struggle against them. The language itself holds the bias, not just the prompt structure.

Meng: From an engineering viewpoint, this means if we are deploying these models regionally, we need to design custom intervention strategies for each language family instead of applying a global patch. A uniform fix won't work here.

Lalam: We can’t just apply a universal fix when the culture itself is where the bias lives; we must create locally aware debiasing protocols to improve our cultural output and ensure representation.

Jane: It’s clear the effectiveness of these prompts is highly dependent on the regional context, which adds another layer to this complex problem that requires our attention.

Tom: And it’s not just about English-centric fixes; the researchers found that while complex prompts help in some Dravidian contexts, they simply don't work as effectively in Indo-Aryan languages. That was a major gap we need to address.

Conclusion & Wrap-Up: Jane: So, we've seen how they measure the bias and how they try and fix it; now let’s look at the bigger picture and wrap up our discussion of "Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations."

Tom: We have a powerful message here: AI is not neutral when cultural biases are involved. It’s not a neutral tool; it’s a mirror of society, and sometimes society is biased, which the LLM reflects.

Lu: The fact that they found higher bias scores for married and single Hindu women in Indo-Aryan to-do lists is such a clear illustration of how specific cultural norms are being reinforced by the data we feed them.

Meng: I think this research forces us to think about the limitations of current models—that self-debiasing isn't enough, especially in regions where cultural norms like purdah are so strongly encoded. We need more than just a prompt suggestion.

Lalam: We need to ensure that our future AI development is not just technically proficient, but culturally sensitive, reflecting the full spectrum of human experience in diverse communities across South Asia.

Jane: It’s a crucial call for fairness and for acknowledging the complexity of marginalized groups in AI output; we can't ignore these systemic issues anymore.

Tom: Absolutely, Jane. We've seen how bias manifests across stories and to-do lists, and we've seen that debiasing methods aren't a magic fix for deeply embedded sociocultural biases.

Lu: I think the possibilities are huge; we can create an AI that truly respects the diversity of South Asia by applying these specific, localized fixes to acknowledge cultural depth.

Meng: This paves the way for developing more robust, culturally aware debiasing frameworks tailored to regional linguistic and social contexts instead of a single global standard.

Lalam: This paper has shown us a path toward improving cultural representation in AI by mitigating those subtle biases that are reinforcing harm in our digital lives.

Tom: That is a powerful message to end on. Thank you all for sharing your insights into "Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations."

Conclusion: Tom: So, to wrap up, we’ve seen a really detailed look at how ingrained cultural norms manifest as measurable biases in AI outputs.

Jane: It has been a deep dive into some incredibly complex and important ground. I think what stands out most is that this research makes it impossible to treat AI generation as anything other than a reflection of the data it consumes.

Lu: I’m just so energized by the potential here; we are talking about building models that don't just process language, but understand cultural nuance with genuine respect.

Meng: From an engineering standpoint, that potential comes with massive requirements for transparency and rigorous testing across every regional dataset, which is crucial if these systems are to be reliable.

Lalam: Exactly. It moves the conversation from abstract ethics into concrete development protocols—how do we actually embed dignity into the code itself?

Jane: It really is a call to action for all of us in this field: we must design for fairness from day one, not as an afterthought patch.

Tom: And it’s clear that a one-size-fits-all approach simply won't cut it; the localized effort suggested by studying *Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations* is the way forward.

Lu: It gives us such a powerful framework for thinking about cultural representation across global deployments.

Meng: We need to build those scalable, region-specific frameworks that can handle this level of complexity consistently.

Lalam: Ultimately, it’s about creating an equitable digital mirror that shows the world in all its beautiful diversity.

Jane: Thank you so much to everyone for joining us on this journey through bias mitigation; it's a vital discussion for the future of AI.

Tom: Indeed, and with that thoughtful conclusion, we'll leave you with these challenges in mind as we pivot to discussing how multimodal models are starting to tackle similar representational issues across different sensory inputs.

Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, Ziwei Zhu, George Mason University

cs.CL

Submitted: 2026-05-06

Updated: 2026-08-25

Code: https://github.com/mamnuya/purdah_and_

Importance score: 75/100

The gist: " Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia.

Key concepts

Intersectionality
The researchers use this concept to analyze how bias is not a single factor but a combination of identity markers. By combining traits like gender and marital status, they uncover specific, often harmful stereotypes that standard evaluation metrics fail to detect.
Language Family Differences
The hosts discuss how debiasing solutions are not one-size-fits-all. Simple prompts have limited power, especially in Indo-Aryan texts, requiring complex, culturally specific intervention strategies for each language family.

Terminology

Summary

"

Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. This study addresses these gaps by investigating prevalent South Asian biases—which include gender, religion, marital expectations, childbearing expectations, and practices like patriarchy and purdah—that are frequently reinforced in generative tasks. The authors note that existing research has challenges including overlooking linguistic and cultural diversity for Indo-Aryan and Dravidian languages in South Asia, neglecting intersectional factors such as childbearing and marital status, or providing limited insight into intersectional bias in open-ended generation.

The researchers propose a novel framework to analyze these culturally specific biases. Key methodological steps include:

  1. Constructing a Bias Lexicon: The authors developed the first intersectional bias lexicon for South Asia, which captures terms related to purdah and patriarchy across various identity dimensions (gender, religion, marital status, and number of children). This lexicon includes 923 culturally grounded terms derived from extensive literature.

  2. Designing Applications: To capture implicit biases in everyday use cases, the study employed three open-ended applications: daily to-do lists, descriptions of hobbies and values, and storytelling.

  3. Generating Data: The researchers created the first large-scale dataset consisting of 100,800 entries. This dataset spans 10 languages (6 Indo-Aryan: Bengali/Bangla, Hindi, Urdu, Punjabi, Marathi, Gujarati; 4 Dravidian: Telugu, Kannada, Malayalam, Tamil), covering 48 identity combinations across the four dimensions.

  4. Evaluating Bias: The study utilizes a lexicon-based metric called Bias TF-IDF to quantify bias. This method calculates the BiasScorei,a,m by summing the Bias TF-IDF values of all matched terms within a specific application (a) and prompting method (m).

  5. Applying Debiasing Strategies: The researchers tested two strategies: a Simple Debiasing prompt (a general instruction to remove bias) and a Complex Debiasing prompt (specific instructions to remove intersectional bias by identity dimensions).

The analysis reveals significant cultural stigmas and biases in the LLM outputs:

  • Cultural Stigmas: The dataset reveals cultural stigmas linked to purdah and patriarchy in Indo-Aryan regions correlated with higher bias levels.

  • Bias Manifestation: The study found that the most bias [is found] in task-oriented generations, especially to-do lists generations.

  • Gender Bias: Figure 5 confirms that females face more bias, particularly in the Indo-Aryan to-do list outputs, where the largest gender gap was observed (female: 0.307 vs. male: 0.043).

  • Religious Bias: Figure 6 shows a major finding that Hindu identities show higher average bias scores than Muslim ones, which the authors note is directly contradicting prior English-language studies.

  • Marital Status Bias: The results indicate that married individuals receive the highest bias scores often via positive associations, while single individuals commonly receive the second-highest scores, suggesting models valorize marriage and stigmatize the unmarried, especially for women.

The study examined how self-debiasing prompts impact these biases:

  • Indo-Aryan Language Failure: The findings show that self-debiasing is largely ineffective in Indo-Aryan texts, where socio-cultural norms like purdah remain encoded. Statistical tests confirmed that "Indo-Aryan languages... yield p > 0.2, indicating that neither simple nor complex debiasing produce statistically significant reductions in bias scores."

  • Dravidian Success: In contrast, Complex prompts significantly reduce bias in Dravidian texts for hobbies/to-do lists (p ≤ 0.02).

  • Overall Limitations: The authors conclude that self-debiasing is insufficient for deeply embedded sociocultural biases, highlighting a need for robust, culturally sensitive multilingual debiasing tactics.

The researchers acknowledge several constraints:

  • Model Limitations: They found that other open-source models (mT5, Aya 101, and Indic-Gemma) were unusable or prohibitively slow, making mT0-xxl the optimal choice.

  • Bias TF-IDF Limitations: The Bias TF-IDF metric itself is limited as it cannot detect contextual or semantic shifts in meaning and may overlook subtle biases that were not recorded in the bias lexicon.

Improvements for AI systems

Based on a rigorous analysis of this research, I propose several critical, specific improvements to current AI systems to mitigate deeply embedded cultural and intersectional biases, particularly in multilingual South Asian contexts.

Improvement: Integrate the novel Bias Lexicon and the Bias TF-IDF (Term Frequency - Inverse Document Frequency) framework into standard evaluation pipelines for LLMs operating in Indo-Aryan and Dravidian languages. This moves beyond simple metrics like toxicity or general gender bias.

Improved System Functionality:

  • The system will automatically detect and quantify subtle, culturally specific stigmas that are invisible to existing benchmarks (e.g., identifying barren associations for childless individuals, or the link between housewife and specific religious/marital statuses).

  • It will provide a granular Bias Score (BiasScore i,a,m), allowing researchers to pinpoint exactly where and how bias manifests across different identities (i), applications (a, e.g., To-do List vs. Story), and prompting methods (m).

Sources

Related papers