Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

arXiv:2504.06313 · cs.CV, cs.CY · Submitted 2025-04-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities".

Tom: This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Now that we've looked at the findings from "Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities," we need to wrap up by thinking about what this research actually means for us as a whole. The authors, Abdulkareem Alsudais, have essentially mapped out how these powerful text-to-image models are drawing on existing cultural biases in their image generation process.

Jane: It’s true; the paper does a lot of work detailing the specific statistical correlations they found between attire and geography or income groups across various activities. The implication is that when we use these tools, especially DALL-E three and Gemini three Pro Preview, we have to be aware that the visual output might not be a neutral reflection of reality but rather a reinforcement of established cultural stereotypes <ref:2504.06313#pg0,DALL-E 3 and Gemini 3 Pro Preview>.

Lu: I see this as an invitation for researchers across the board to develop better evaluation metrics that specifically target these kinds of cultural and geographical elements, which is something the paper hints at by mentioning prior work in this area.

Meng: From my side, I think the practical implication is that we need robust methods for auditing these models to ensure they are not unintentionally promoting harmful visual tropes in real-world deployments where users might interact with them.

Lalam: I believe this study gives us concrete data points—like the twenty-eight point four percent figure or the associations with MENA and income groups—that move this conversation past just being theoretical and into actionable territory for improving how AI is trained for better cultural depiction <ref:2504.06313#pg0>.

Tom: So, to put it simply, the title of this paper highlights that we are studying how these text-to-image models represent people from different nationalities when asked to generate images of common activities, and the authors conclude that these models frequently depict people wearing traditional or impractical clothing in ways that statistically correlate with specific regions and income groups.

Jane: And what this means for us is a reminder that as AI gets more integrated into daily life, we need to critically examine the visual outputs to ensure they aren't just repeating old, biased imagery without prompting us to question those images.

Lu: This work suggests that future research should focus on creating methods that can actively identify and correct these representation patterns within the models themselves rather than just reacting after the fact.

Meng: And I think we need engineers to prioritize building tools for bias detection because if we don't, these visual biases will just get cemented into our systems.

Lalam: Ultimately, this paper shows that understanding these visual patterns can be a tool for making the AI culture more inclusive by showing us exactly where the representation is skewed and how we can work to correct those specific imbalances.

Conclusion: Tom: So we’ve seen how DALL-E three and Gemini three Pro Preview create images of people from different backgrounds doing everyday things, and now we’re getting to the conclusion of this study by Abdulkareem Alsudais and his team.

Jane: Right, so to recap, the paper looked at how these models show people across two hundred six nationalities when asked for pictures of common activities like cooking or jogging. The main takeaway is that these models often default to showing traditional or sometimes just impractical clothing, and this tendency isn't random; it’s tied directly to a person's region and income level.

Lu: It’s fascinating because it shows that the AI isn't just guessing; there are real statistical patterns emerging in how cultural representation is being baked into these systems, which opens up incredible avenues for exploring how these models learn societal biases.

Meng: From an engineering standpoint, seeing those strong correlations between attire and income groups tells us exactly where we need to focus our data curation efforts if we want to build more representative AI tools. It shows the model is reflecting the data it was trained on in a very specific way.

Lalam: I think this is hugely important because by quantifying these patterns, we gain a roadmap for actively de-biasing future generations of image models so they don't just repeat old visual stereotypes. This research gives us concrete evidence to fight that visual bias in the AI ecosystem.

Tom: Exactly! So, when we look at the title, "Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities," it really sums up this deep dive into how these tools capture cultural context through clothing choices.

Jane: It really does; it’s not just about making pretty pictures; it’s about understanding the underlying assumptions those models make about who is doing what and how they look while doing it. This points to a deeper societal issue reflected in technology.

Lu: I see this as a huge opportunity for creative AI exploration, because if we can map these patterns so clearly, we can intentionally steer the models toward more nuanced or diverse outputs in future iterations. It’s like mapping the uncharted territories of visual representation.

Meng: I think the real impact here is on deployment; if we understand these income and regional correlations, our teams at least know which datasets need to be prioritized for balancing before we roll out new applications. That’s practical application right there.

Lalam: For me, the most impactful vision here is that this detailed analysis provides the necessary framework for building AI systems that can intentionally promote more diverse and respectful cultural depictions in all forms of visual media. It gives us the tools to build a better digital culture through better AI design.

Tom: So we’ve covered how these models lean into certain visual tropes based on where a person is from or what they do, and it really shows the power of this research to make our AI systems more aware of their own cultural fingerprints.

Prince Sattam bin Abdulaziz University

cs.CV, cs.CY

Submitted: 2025-04-08

Updated: 2026-04-11

Journal ref: Journal of Artificial Intelligence Research, Vol. 87, Article 8. Publication date: September 2026

DOI: 10.1613/jair.1.22039

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in

Key concepts

Text-to-Image (T2I) Models
These are AI systems that create images based on written descriptions, like prompts. The researchers tested DALL-E 3 and Gemini 3 Pro Preview to see how well they represent people from different countries when asked to show them doing normal things.
Representational Patterns
This refers to the recurring ways the AI models choose to depict people in their generated images. The study found that certain clothing styles, like traditional or impractical outfits, appear more often than others across different activities and nationalities.
Statistical Associations
This involves using math to see if patterns observed in the images (like wearing traditional clothes) are truly related to other factors, such as a person's country of origin or their income group. The study found significant links between these visual patterns and demographic data.
Alignment Scores
These scores measure how well an image actually matches the text prompt used to create it. The research discovered that when using tools like CLIP, ALIGN, and GPT-4.1 mini, the AI gave higher alignment scores to images featuring traditional clothing.

Terminology

Summary

This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompted to generate images of individuals engaging in common everyday activities. The key finding is that these models frequently depict individuals wearing traditional attire or impractical clothing for the specified activities, with these patterns being statistically significantly associated with specific regions and income groups.

Research Objectives

The main objective of the paper is to study how T2I models represent individuals from various countries when prompted to generate images of people engaging in typical activities. Specifically, the paper focuses on three primary research questions:

** RQ1: When generating images of people from various countries performing common activities, to what extent are they depicted in traditional or impractical attire, and how does this correlate with the country's region or income group?**

** RQ2: When employing CLIP, ALIGN, and GPT-4.1 mini to calculate alignment between images and prompts, to what extent do alignment scores correlate with the presence of traditional or impractical attire in labeled images?**

** RQ3: When examining revised prompts, what elements do they introduce to the final instructions, and how do these elements contribute to the representation patterns identified in RQ1?**

Methodology and Dataset Construction

The study constructs a dataset that includes 206 nationalities across five scenarios: couples cooking at home, groups playing soccer in public parks, a person working on their laptop from home, a group jogging, and a group drinking coffee at a coffeeshop. Images were generated using DALL-E 3 (OpenAI) and Gemini 3 Pro Preview (Google DeepMind). The dataset includes 206 images per model across the five scenarios, totaling 2,060 images. The images are manually labeled to indicate whether the depicted individuals are dressed in traditional or impractical attire.

Analysis of Representational Patterns (RQ1)

When aggregating results across all activities and models, 28.4% of the images depicted individuals wearing traditional attire. Among the subset of images assessed for impractical clothing in the two athletics-related activities, 15.9% depicted individuals wearing attire that was impractical for those activities. The prevalence of traditional attire varied by activity; "the cooking and coffee activities yielding the highest percentages of images labeled as “traditional,” whereas the working activity yielded the lowest percentage, likely because it specified a single individual in the prompt."

Statistical Associations with Region and Income Group

The analysis stratified results by region and income groups to identify systematic patterns. The regional analysis revealed that "the Middle East & North Africa region consistently exhibited the highest percentages across activities and label categories. For the 14 activity–label combinations, Middle East & North Africa (MENA) ranked first in the percentage of labeled images in 10 cases and second in the remaining four cases, behind South Asia. Across income groups, Low Income countries exhibited the highest percentages in most scenarios, often followed by Lower Middle Income countries. When combining all 10 sets of results across both models and five activities, 28.4% of images depicted people dressed in traditional attire and strong associations were observed with both region (χ2 = 386.1, p <.001, Cramér’s V = 0.43) and income group (χ2 = 171.8, p <.001, Cramér’s V = 0.28)."

Evaluation Model Efficacy (RQ2)

The second research question examined the effectiveness of CLIP, ALIGN, and GPT-4.1 mini for evaluating alignment between prompts and generated images. For the standard prompts used to generate the images, an unexpected pattern emerged: Across all three scoring methods, alignment scores were consistently higher for images featuring individuals in traditional clothing compared to those without it. This pattern was also held when aggregating results across all activities for both DALL-E 3 and Gemini 3 Pro Preview. For example, for DALL-E 3, the average scores were 31.31 vs. 32.28 (CLIP), 23.11 vs. 25.36 (ALIGN), and 67.63 vs. 75.70 (GPT-4.1 mini).

Analysis of Revised Prompts (RQ3)

The final research question examined the revised prompts generated by OpenAI for the standard prompts across countries, focusing only on DALL-E 3 images. An initial exploratory analysis processed all revised prompts to identify frequently used words, noting that the revised prompts also introduced words not explicitly present in the original prompts, including “traditional” and “diverse.” Specifically, all revised prompts for each activity were processed to count how often the word “traditional” was added, finding it appeared in "88.

Improvements for AI systems

Here are specific, actionable improvements to AI systems derived from the findings of this research:


) 1. Implement a Multi-Stage Bias Detection and Mitigation Pipeline in T2I Generation:

The paper demonstrates that representational biases (traditional attire, impractical clothing) are introduced across the entire pipeline—generation, evaluation, and prompt revision.

  • [System Improvement]: Integrate a pre-generation bias check layer using metrics derived from RQ1 (labeling criteria) before final output. This layer should flag prompts or generation seeds likely to produce stereotypical outputs based on known regional/income group correlations (MENA, Sub-Saharan Africa, Low Income).

  • [System Improvement]: Develop a Prompt Sanitization module that analyzes the generated revised prompts (RQ3 analysis) for the frequent insertion of culturally essentializing terms like traditional, especially when combined with activity prompts that are contextually inappropriate. If such terms appear frequently without explicit user intent, the system should automatically rewrite them to be more context-appropriate or remove them entirely, guided by a learned preference from the non-traditional image subset analysis (Table 8).

  • [System Improvement]: Incorporate evaluation metrics that go beyond simple CLIP similarity (RQ2). Develop a Contextual Validity Score that specifically penalizes high alignment scores for images labeled as both traditional and impractical in athletics, shifting the reward function away from simply matching the text to capturing contextual appropriateness.

  • [Improved AI System Capability]: The system will move from merely generating images based on text to actively self-correcting its outputs by identifying and neutralizing biased cues introduced during the generation or refinement stages. It will ensure that generated imagery for common activities (e.g., soccer, jogging) adheres strictly to functional appropriateness, regardless of underlying biases in the training data or prompt revision mechanisms.

) 2. Develop a Context-Aware Evaluation Framework for Vision-Language Models (VLMs):

The study shows that standard alignment scores (CLIP, ALIGN) can be misleading because they prioritize nationality cues over activity relevance, leading to systematic misjudgments of stereotypical depictions.

  • [System Improvement]: Replace reliance on single alignment metrics with a weighted scoring system that dynamically adjusts the importance of different text elements based on the activity context. For example, when the prompt specifies playing soccer, the weight given to terms related to athletic attire should be prioritized over terms related to cultural identity (country/nationality).

  • [System Improvement]: Implement a comparative evaluation mechanism using multiple models (DALL-E 3 vs. Gemini 3 Pro) alongside different scoring methods. If one model consistently assigns higher alignment scores to stereotypical representations, the system should flag that model's output for human review or apply a learned bias correction factor specific to that model's known tendencies (e.g., GPT-4.1 mini's tendency in RQ2).

  • [System Improvement]: Create an explicit Construct Validity Check in the evaluation phase, using the findings from RQ2 and Table 6/7. This check will compare alignment scores for traditional vs. non-traditional groups across all activities, ensuring that high alignment is not solely driven by the presence of nationality cues when those cues are contextually irrelevant.

  • [Improved AI System Capability]: The system will provide a more trustworthy evaluation layer that judges an image based on its functional relevance to the specified activity, rather than being unduly influenced by learned correlations between nationality and cultural attire embedded in the evaluation models themselves. This leads to a more robust assessment of representational fairness across diverse user prompts.

) 3. Create Dynamically Contextualized Prompt Engineering Strategies:

The analysis of revised prompts (RQ3) revealed that the inclusion of terms like traditional is heavily correlated with higher traditional labels, suggesting prompt rewriting can inadvertently reinforce stereotypes.

  • [System Improvement]: Implement a Prompt Revision Strategy that uses semantic clustering (as done in RQ3) to identify and suppress high-frequency, culturally essentializing vocabulary (traditional, specific regional descriptors) when the activity context does not necessitate them.

  • [System Improvement]: Develop De-Stereotyping Prompts as an automated countermeasure. If the system detects a high probability of generating an impractical or stereotypical image (e.g., based on prior results), it should automatically generate a prompt variation that explicitly requests contextually appropriate attire, such as: Generate a realistic photograph-like image of a group from [country] playing soccer in a public park, ensuring all individuals are wearing appropriate athletic wear.

  • [System Improvement]: Use the thematic clustering (Table 9) to create activity-specific prompt templates. For instance, the Soccer cluster vocabulary should automatically trigger constraints on attire when generating images for that activity, overriding potentially biased general language learned from other activities.

  • [Improved AI System Capability]: The system will evolve beyond simple text-to-image generation by employing intelligent prompt engineering that proactively steers the model away from known representational pitfalls, ensuring that user intent (e.g., play soccer) is translated into a depiction of contextually appropriate action, rather than defaulting to culturally exotic tropes.

Sources

Related papers