CultureScore: Evaluating Cultural Faithfulness in Video Generation Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "CultureScore: Evaluating Cultural Faithfulness in Video Generation Models".
Tom: CultureScore is a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions—Identity, Context, and Behavior—to diagnose where current video generation models diverge from authentic cultural representation.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, to recap, we’re talking about the "CultureScore: Evaluating Cultural Faithfulness in Video Generation Models" paper. The title itself sets a clear goal for this research, which is to move past just looking at visual quality <ref:2606.07311#pg0>.
Jane: It's about creating a new way to measure how well video generation models capture cultural representation across different areas like identity and behavior <ref:2606.07311#pg1>. The authors are Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes, and Mahdi M. Kalayeh <ref:2606.07311#pg0>.
Lu: It’s interesting to see the collaboration across institutions; MIT and Mila – Quebec AI Institute are involved in this study, which suggests a deep dive into foundational research <ref:2606.07311#pg1>.
Meng: I'm curious about the context of this work. Are these models typically trained on broad datasets that might lead to these kind of superficial understandings we just discussed?
Lalam: The authors explicitly state that current metrics like VideoScore fail because they don’t have a mechanism to assess cultural faithfulness, which is what this paper is addressing <ref:2606.07311#pg0>.
Tom: That’s the core problem: we have models that look great but might be culturally wrong, and CultureScore provides a diagnostic tool for that misalignment <ref:2606.07311#pg1>.
Jane: It’s about giving us a more granular way to look at where the AI diverges from authentic representation through Identity, Context, and Behavior <ref:2606.07311#pg1>.
Lu: Think of it like this: instead of just a big score on quality, we get detailed feedback on whether the person in the video is dressed correctly for that setting or if their interaction looks culturally appropriate <ref:2606.07311#pg0>.
Meng: That specificity would be incredibly useful for our engineering teams because it tells us precisely which part of the generation pipeline is failing, not just the output itself.
Lalam: Because they operationalize this by building an evaluation suite with two thousand nine hundred forty-three culturally validated prompts across ten countries and five socio-cultural domains <ref:2606.07311#pg1>.
Tom: That’s a big dataset to work with for testing these three dimensions—Identity, Context, and Behavior—across so many different cultural scenarios <ref:2606.07311#pg1>.
Jane: And those QA pairs they generated across identity, behavior, and context are what allow them to produce those component-level accuracy scores that aggregate into the final CultureScore <ref:2606.07311#pg1>.
Lu: The paper is really setting a new standard for how we evaluate these complex multimodal outputs in terms of cultural representation <ref:2606.07311#pg2>.
The paper's summary: Tom: Now that we know what it’s called and who did it, let’s get into the actual substance of the paper. Essentially, this paper lays out the CultureScore framework as a way to systematically decompose cultural faithfulness <ref:2606.07311#pg0>.
Jane: It breaks things down into three specific dimensions: Identity—who is represented and how—Behavior—the normative gestures and expressivity—and Context, which covers the culturally situated settings and social conventions <ref:2606.07311#pg1>.
Lu: I think the key insight here is that this compositional approach lets us see subtle cultural mismatches, like an incorrect greeting gesture being performed in a wrong setting <ref:2606.07311#pg0>.
Meng: So, it’s not just a holistic check; it’s a fine-grained metric that exposes those tiny cultural errors that big metrics always miss <ref:2606.07311#pg0>.
Lalam: That granularity is powerful because it moves the evaluation from a simple pass/fail to a detailed diagnostic report on specific cultural components <ref:2606.07311#pg1>.
Tom: And the results they found were pretty intense. The main conclusion is that no current model actually achieves culturally faithful video generation <ref:2606.07311#pg0>.
Jane: They quantified this by showing that the best model reached only a fifty-six point eight percent overall CultureScore <ref:2606.07311#pg0>, with Behavior being the most difficult aspect, staying below fifty-two point one percent across all models <ref:2606.07311#pg0>.
Lu: That fifty-six point eight percent figure is a strong indicator that we still have a significant gap to bridge before we consider these tools truly faithful for cultural representation <ref:2606.07311#pg0>.
Meng: So, the implication is that simply focusing on boosting visual quality metrics isn't going to solve the cultural accuracy problem, which is a really important practical takeaway <ref:2606.07311#pg1>.
Lalam: And they showed that models often rely heavily on explicit geographic tokens as cultural triggers rather than having internalized the underlying cultural concepts <ref:2606.07311#pg2>.
The paper's improvements: Tom: Moving on to what the authors suggest to fix these issues, they propose several ways to improve the situation, focusing heavily on prompt engineering and data enrichment <ref:2606.07311#pg2>.
Jane: They suggest using decomposed and culturally explicit prompt guidance because this helps improve CultureScore across all three models and all three dimensions <ref:2606.07311#pg2>.
Lu: The paper found that for Identity, LTX-two benefited the most with an improvement of eighteen point two percent, but Context showed the largest absolute gains across all models <ref:2606.07311#pg2>.
Meng: So, prompt engineering is a strong lever, but they also noted that Behavior is the most resistant dimension to prompt enrichment; no model exceeded fifty-two point one percent on Behavior even with extended prompting <ref:2606.07311#pg2>.
Lalam: That confirms that temporal coherence in motion sequences remains a persistent failure mode that just adding more text isn't going to fix by itself <ref:2606.07311#pg2>.
Tom: It’s interesting how they found that models are relying on explicit geographic tokens as cultural triggers instead of truly understanding the underlying concepts <ref:2606.07311#pg2>.
Jane: They suggest generating "Geographically Constraint Removed Prompts" to test whether models use explicit country names or if they have internalized the actual cultural concepts <ref:2606.07311#pg2>.
Lu: The paper is trying to force the model out of relying on surface-level triggers and into a deeper understanding of the culture itself <ref:2606.07311#pg2>.
Meng: From an engineering standpoint, this means our prompting pipeline needs to incorporate these types of tests to see if we’re just patching symptoms or actually teaching the model the underlying structure <ref:2606.07311#pg2>.
Lalam: And they also mentioned that providing decomposed cultural guidance improves Identity and Context scores, but Behavior remains stubbornly resistant to prompt enrichment <ref:2606.07311#pg2>.
Conclusion: Tom: So, as we wrap up this discussion on "CultureScore: Evaluating Cultural Faithfulness in Video Generation Models," the main implication is that we need a new evaluation framework for generative AI <ref:2606.07311#pg0>.
Jane: It’s a reminder that optimizing for visual quality alone doesn't guarantee cultural accuracy, and we should be looking at tools like CultureScore as a necessary complement to human evaluation <ref:2606.07311#pg1>.
Lu: The authors are proposing this framework so we have an interpretable diagnostic that reveals exactly where models diverge from authentic depiction across Identity, Context, and Behavior <ref:2606.07311#pg0>.
Meng: For practical application, it means our next steps should involve training models specifically on the high-fidelity cultural data derived from those validated prompts <ref:2606.07311#pg2>.
Lalam: And they also mentioned that we should be cautious about using the Identity dimension for physical markers because there's a risk of reinforcing stereotypes if we rely too much on superficial visual features <ref:2606.07311#pg2>.
Tom: It’s a lot to take in, but overall, CultureScore provides us with a reusable foundation for auditing cultural representation in video generation <ref:2606.07311#pg0>.
Jane: That’s right. We learned that while we can get close, no current model is fully faithful culturally, and we need this kind of detailed measurement to guide future development <ref:2606.07311#pg0>.
Lu: I think the real excitement here is how this framework opens up new avenues for research into how AI actually learns cultural concepts rather than just pattern matching tokens <ref:2606.07311#pg2>.
Meng: And I think the practical impact will be in building better guardrails that check for these specific cultural failures before we ship anything out <ref:2606.07311#pg2>.
Lalam: We can’t wait to see how researchers use this framework to push models past that fifty-six point eight percent threshold and actually achieve a higher degree of cultural representation <ref:2606.07311#pg0>.
Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes, Mahdi M. Kalayeh, Paul Pu Liang
Massachusetts Institute of Technology · Mila – Quebec AI Institute
cs.CV, cs.AI
Submitted: 2026-06-05
Updated: 2026-10-04
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: CultureScore is a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions—Identity, Context, and Behavior—to diagnose where current video
Key concepts
- CultureScore
- A framework that measures a video's cultural faithfulness by breaking it down into three parts: Identity (who is shown), Behavior (how people act), and Context (where they are). It provides a fine-grained metric to spot subtle cultural mistakes.
- Identity, Context, Behavior
- These are the three dimensions CultureScore uses to assess cultural representation. Identity covers who is represented; Context covers the setting and social rules; and Behavior covers gestures, speech patterns, and expressivity.
- Counterfactual Prompt Augmentation
- A method used to test model behavior by adding hypothetical or modified text prompts. This helps researchers understand how much a model relies on specific geographic names versus understanding general cultural concepts.
Terminology
Summary
CultureScore is a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions—Identity, Context, and Behavior—to diagnose where current video generation models diverge from authentic cultural representation. This research matters because existing metrics like VideoScore fail to assess cultural accuracy, meaning a model can achieve high visual quality while fundamentally misrepresenting cultural norms. The study reveals that no current model achieves culturally faithful video generation, with the best-performing model reaching only 56.8% overall CultureScore, underscoring that cultural faithfulness is an essential criterion for equitable video generation tools.
CultureScore Framework
The proposed framework decomposes a textual prompt into three culturally grounded facets: Identity (who is represented and how), Behavior (culturally normative gestures, prosody, expressivity), and Context (culturally situated settings and social conventions). This compositional approach allows for a fine-grained metric
that exposes subtle cultural mismatches, such as an incorrect greeting gesture performed in the wrong setting. The framework operationalizes this by building an evaluation suite grounded in CulturalFrames, spanning 10 countries and 5 socio-cultural domains, yielding 6,174 generated videos across three state-of-the-art models.
Evaluation Methodology
The evaluation involves a four-step process: (i) Counterfactual prompt augmentation to probe model behavior; (ii) video generation using the augmented prompts; (iii) IBC decomposition to evaluate Identity, Behavior, and Context dimensions; and (iv) VLM-based scoring where Vision-Language Models perform fine-grained, aspect-based question answering. This process produces component-level accuracy scores that aggregate into an overall CultureScore. The evaluation suite includes 2,943 culturally validated prompts spanning 10 countries and 5 socio-cultural domains.
Key Findings on Model Performance
The evaluation reveals several striking findings across the three models (Veo 3.1 Fast, LTX-2, and Wan 2.2):
no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8% overall CultureScore with Behavior the most challenging dimension, which remains below 52.1% across all models.
models that score highest on perceptual metrics systematically score lowest on cultural faithfulness, and vice versa — a divergence that is statistically validated by native human evaluators.
The analysis shows that Context consistently yields the highest accuracy across all models, while Behavior remains the most challenging dimension.
Prompt Sensitivity and Cultural Knowledge
The study investigates whether models rely on explicit geographic identifiers or have internalized cultural concepts as open-world knowledge. The results indicate that models rely heavily on explicit geographic tokens as cultural triggers rather than having internalized the underlying cultural concepts.
Specifically, removing country names from base prompts causes a consistent and substantial drop
in accuracy across all dimensions for LTX-2 and Veo 3.1 Fast, suggesting that cultural understanding remains largely surface-level.
Metric Alignment with Human Preference
A critical finding is the inverse relationship between standard video generation metrics and CultureScore: "LTX-2 (HaCohen et al., 2024) secures the highest VideoScore (avg. 3.34) yet ranks second lowest in CultureScore accuracy (avg. 51.4%), indicating that optimizing for perceptual quality does not translate to cultural faithfulness. Conversely, CultureScore is
directionally aligned with preference, with Wan 2.2 ranking highest on both cultural accuracy and human preference among the models evaluated. The study concludes that
general-purpose video quality metrics are not merely insufficient for measuring cultural faithfulness, but can actively mislead model selection when cultural accuracy is the priority."
Prompt Enrichment Effects
Providing decomposed, culturally explicit prompt guidance improves CultureScore across all three models and all three dimensions. For Identity, LTX-2 benefited most (+18.2%), while Context showed the largest absolute gains across all models.
However, Behavior is described as the most resistant dimension to prompt enrichment,
with no model exceeding 52.1% on Behavior even under extended prompting, confirming that temporally coherent motion sequences remain a persistent failure mode that prompt enrichment alone cannot resolve.
Limitations and Ethics
The research acknowledges limitations, including Asymmetric model coverage
due to API costs for Veo 3.1 Fast and the fact that human evaluation is limited by geography (excluding Iran). Furthermore, the reliance on Qwen3-VL for scoring introduces a risk of VLM-as-judge biases.
The authors recommend treating CultureScore as a complementary metric to be used alongside human evaluation, rather than as a standalone ground truth,
especially concerning the Identity dimension which risks reinforcing stereotypes.
Conclusion
The work proposes CultureScore as a reusable foundation for auditing cultural representation in generative AI by providing an interpretable diagnostic that reveals where and how models diverge from authentic cultural depiction.
Improvements for AI systems
Here are specific, actionable improvements for AI video generation systems based on the CultureScore framework:
-
The core improvement is transitioning from holistic quality metrics (like VideoScore) to a compositional, three-dimensional cultural evaluation framework (CultureScore). This allows for diagnostic failure analysis rather than just a single score.
-
The improved system can be used to generate an
Interpretability Report
for any video output, detailing exactly where it succeeded or failed culturally:
Narrow the diagnosis to three granular dimensions:
-
[Identity]: Accuracy of physical appearance, attire, and demographic markers (e.g., correctly depicting a specific national garment or regional clothing).
-
[Behavior]: Fidelity of culturally normative gestures and interactions (e.g., distinguishing between a generic handshake and a culturally specific greeting like the Namaste or Salam).
-
[Context]: Accuracy of culturally localized background details, including social conventions, environmental decor (e.g., correct placement of traditional rugs or furniture), and social arrangements.
- Implement prompt engineering pipelines that actively test for cultural reliance versus internalization:
-
Generate
Geographically Constraint Removed Prompts
to test if models rely on explicit country names (triggers) or have internalized the underlying cultural concepts (open-world knowledge). -
Use
Extended Prompting
grounded in dictionary definitions of behaviors and contexts to force the model to synthesize complex, nuanced visual descriptions.
- Develop a proactive feedback loop for model training:
-
Train models specifically on high-fidelity cultural data derived from the 2,943 culturally validated prompts.
-
Use
Question Generators
(as described in Section C.4) to create a massive, automatically generated QA dataset of thousands of component-level accuracy scores. This allows for fine-tuning the model's reward function to prioritize maximizing CultureScore dimensions over simple perceptual metrics like VideoScore, thus directly addressing the observed inverse correlation where high visual quality models score poorly on cultural faithfulness.
- Integrate a VLM-based scoring mechanism (like Qwen3-VL) as an
AI Cultural Auditor
:
- This auditor should be deployed post-generation to automatically query the video against the decomposed CultureScore questions, providing component scores (Identity: X%, Behavior: Y%, Context: Z%). This provides a quantitative, aspect-based metric that is more reliable than holistic scores.
- Incorporate human preference alignment into model selection criteria:
- Use the findings from Section 4.1 (Figure 8) to establish a
Cultural Faithfulness Threshold.
Models should not only be selected based on high VideoScore but must also demonstrate a favorable correlation with human preference rankings, as the paper showed that perceptual quality metrics actively mislead selection when cultural accuracy is the priority.
- Mitigate Stereotyping Risks:
- Implement rigorous checks during training and inference to ensure that the evaluation of Identity dimension (physical markers) does not reinforce harmful stereotypes by grounding it in documented, diverse cultural atlases rather than relying solely on superficial visual features.
Abstract
As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that replaces a Namaste with a handshake receives the same score as one that generates the gesture correctly. We propose CultureScore, a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions: Identity (who is represented), Context (culturally localized background), and Behavior (normative gestures and interactions). We operationalize this framework through an evaluation suite spanning 10 countries, yielding 6,174 generated videos across three state-of-the-art models. Our evaluation reveals that no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8% overall CultureScore, with Behavior the most challenging dimension; no model exceeds 52.1% on behavior. Furthermore, the highest-scoring model (LTX-2) on visual quality was ranked last by native annotators, while CultureScore's Behavior dimension shows the strongest positive correlation with human cultural judgment among the automatic metrics we evaluate, underscoring that cultural faithfulness is an essential criterion for equitable video generation. Data and code are publicly available.
Sources
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive?
- CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
- LTX-Video: Realtime Video Latent Diffusion
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
- Scalable Diffusion Models with Transformers
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- Wan: Open and Advanced Large-Scale Video Generative Models
- Unified Reward Model for Multimodal Understanding and Generation
- Qwen3 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models