On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

arXiv:2608.11002 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Sicheng Zhang, Zhonghao Yan, Binzhu Xie, Shi Qiu, Muzammal Naseer, Naveed Akhtar, Mubarak Shah

Khalifa University · Queen Mary University of London · The Chinese University of Hong Kong · The University of Western Australia · The University of Melbourne · University of Central Florida

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted to ACM MM 2026

Code: https://github.com/RISys-Lab/LingT2I

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: The paper introduces LingT2I, a new benchmark designed to evaluate cross-lingual effects in text-to-image (T2I) generation.

Terminology

Summary

The paper introduces LingT2I, a new benchmark designed to evaluate cross-lingual effects in text-to-image (T2I) generation. The benchmark covers 10 widely used languages (English, Chinese, Hindi, Spanish, Arabic, French, Portuguese, Russian, Japanese, and Korean) with 33K prompts, split into two tasks: Content Generation (30K prompts) and Text Rendering (3K prompts).

The authors evaluate 17 recent T2I models on this benchmark, including general-purpose models (e.g., SD3.5, SDXL, FLUX, Qwen-Image, Nano Banana), multilingual-enhanced models (e.g., PEA, X2I, MuLan), and rendering-oriented models (e.g., AnyText, AnyText2, EasyText). Their comprehensive analysis reveals several key findings:

1. Linguistic Inequality: General-purpose models exhibit severe linguistic inequality, with performance skewed toward high-resource Indo-European languages. For example, in the Content Generation task, English achieves an average CLIPScore of 0.78, while Hindi and Arabic only reach 0.38. Multilingual-enhanced models show lower variance but often at the cost of reduced performance in privileged languages. The paper states: multilingual-enhanced models exhibit much lower variance, suggesting more balanced cross-lingual performance, while most general-purpose models suffer from severe linguistic inequality.

2. Text Rendering Failures: Non-Latin scripts (e.g., Arabic, Hindi, Korean) often appear broken, unreadable, or hallucinated. The paper notes: non-Latin writing systems remain a major bottleneck, leading to broken or unreadable text rendering. Even specialized rendering models struggle, with EasyText achieving the best balance (average precision 0.67, standard deviation 0.14) but still showing pronounced disparities.

3. Language-dependent Trade-offs: For semantically identical prompts, different languages exhibit divergent trade-offs across evaluation dimensions (e.g., realism, semantic faithfulness, style). The paper states: "even for semantically identical prompts, different languages exhibit divergent trade-offs across dimensions such as realism, semantic faithfulness, and style; these interactions may appear as coupled improvements, conflicting trends, or balanced compromises." For example, in the Toxicity-Style trade-off, Hindi and Arabic produce safer but less stylistically consistent images, while English emphasizes coherent aesthetics at the cost of higher cultural bias.

4. Language-dependent Generation Patterns: The paper reveals systematic variations in generation behavior across languages, including:

  • Demographic Bias: Strong language-demographic alignment, where Hindi prompts yield predominantly Indian subjects (94.6%), Japanese/Korean prompts produce over 65% Asian subjects, and Western languages (Russian, French, English) show 84.9%–97.1% White-presenting subjects.

  • Cultural Tendency: 23.2% of generated images contain identifiable culture-specific visual elements, with 79.6% of those matching the prompt language's cultural region. The paper notes: multilingual prompting steers generation toward language-specific cultural aesthetics, expressed through cues such as writing systems, architecture, clothing, and symbolic objects.

  • Rendering Errors: Different failure modes for Latin (stable prefix, fragile suffix pattern) versus CJK scripts (structural collapse of sequence integrity).

5. Causal Analysis: The paper identifies three failure patterns in text rendering: semantic substitution, script-specific structural failure, and script-selection/romanization bias. It also finds that transliteration does not improve content alignment (scores dropped from 0.78 to 0.49 on average), and that prompt fragmentation shows a strong negative correlation with generation quality (Spearman rank correlation rho = −0.89), identifying inefficient tokenization as a systematic bottleneck for non-Latin languages.

The paper's key contributions are:

  • Introducing LingT2I, a new dataset covering 10 languages with 33K prompts for analyzing cross-lingual effects in both Content Generation and Text Rendering.

  • Providing the first comprehensive cross-lingual analysis revealing linguistic inequality and language-specific trade-offs across dimensions.

  • Revealing various language-dependent generative patterns providing insights for model design.

The authors conclude with insights for future research: "multilingual capability should be achieved through native architectural design rather than post-hoc adaptation; training data should be organized by language family and curated with cultural grounding; models should maintain balanced performance across evaluation dimensions; bias-aware data curation and translation-based augmentation may help mitigate cultural and demographic biases; models could benefit from script-specific rendering modules or training strategies tailored to their structural characteristics."

Improvements for AI systems

Improvements to AI Systems Based on LingT2I:

  1. Implement language-family-aware tokenization
  • Replace generic subword tokenizers with script-specific tokenizers (e.g., separate handling for Hangul syllable blocks, Devanagari conjuncts, Arabic diacritics).

  • Add a tokenization quality monitor that flags high fragmentation rates (ρ = −0.89 correlation with quality) and triggers fallback to character-level or syllable-level encoding.

  • Resulting capability: Non-Latin prompts will generate coherent images with fewer structural collapses, especially for CJK and Indic scripts.

  1. Add a cross-lingual semantic alignment layer
  • Train a multilingual encoder (e.g., XLM-R) to map prompts from 10 languages into a shared semantic space before feeding into the T2I diffusion backbone.

  • Use contrastive learning on semantically identical prompts (from LingT2I) to reduce CLIPScore variance across languages (target: reduce English-to-Hindi gap from 0.40 to <0.15).

  • Resulting capability: Users can prompt in Hindi or Arabic and get image quality comparable to English, without needing translation.

  1. Introduce script-aware text rendering modules
  • Add a dedicated rendering head that switches between Latin, Arabic, Devanagari, Hangul, and CJK glyph generators based on detected script.

  • Train this head on the 3K Text Rendering prompts with script-specific loss functions (e.g., stroke-order loss for Hangul, ligature preservation for Arabic).

  • Resulting capability: Generated images will contain readable, correctly-formed text in non-Latin scripts, reducing broken or hallucinated characters.

  1. Deploy a cultural grounding filter
  • Post-process generated images with a culture classifier (trained on LingT2I’s cultural tendency annotations) to detect mismatched cultural cues (e.g., Indian architecture for a Japanese prompt).

  • Apply targeted diffusion refinement (e.g., inpainting) to replace mismatched elements with culturally appropriate ones from a curated visual dictionary.

  • Resulting capability: Images will match the prompt’s cultural region (e.g., 79.6% accuracy → >90%) while reducing demographic bias (e.g., Hindi prompts no longer default to 94.6% Indian subjects unless explicitly requested).

  1. Build a trade-off-aware generation controller
  • Train a meta-model that predicts, for a given language and prompt, the optimal balance between realism, semantic faithfulness, and style (based on LingT2I’s trade-off patterns).

  • Dynamically adjust diffusion guidance weights (e.g., CFG scale, style injection strength) per language and dimension.

  • Resulting capability: Users get consistent quality across languages—e.g., Hindi prompts will no longer sacrifice style for safety, and English prompts will not over-prioritize aesthetics at the cost of cultural bias.

  1. Add a transliteration-aware prompt optimizer
  • Detect when a prompt is in a low-resource script and automatically generate a hybrid prompt (original script + romanized/transliterated version) for the text encoder.

  • Use a learned weighting to combine embeddings, avoiding the 37% drop in CLIPScore observed with naive transliteration.

  • Resulting capability: Users can input native scripts, and the system will internally boost alignment without degrading content fidelity.

  1. Integrate a linguistic-inequality diagnostic tool
  • Add a runtime evaluator that measures CLIPScore variance and text-rendering precision across the 10 LingT2I languages for any new prompt.

  • If variance exceeds a threshold, automatically trigger fine-tuning on underrepresented language families (e.g., using LoRA on Hindi/Arabic data).

  • Resulting capability: The system self-corrects over time, reducing linguistic inequality in deployed models without manual intervention.

What the improved AI system can do:

  • Generate images from prompts in 10 languages with near-equal quality (CLIPScore variance < 0.1 across languages).

  • Render readable, grammatically correct text in Arabic, Hindi, Korean, and Japanese within images.

  • Produce culturally appropriate visuals that match the prompt’s language region, with reduced demographic stereotyping.

  • Automatically balance realism, style, and safety per language, avoiding language-specific trade-off failures.

  • Self-diagnose and mitigate cross-lingual performance gaps during deployment.

Abstract

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely used languages with 33K prompts, designed to evaluate cross-lingual effects in both content generation and text rendering. Building on this benchmark, we conduct a comprehensive cross-lingual analysis, uncovering linguistic inequality and language-dependent trade-offs across evaluation dimensions. Beyond quantitative evaluation, we further reveal a range of language-dependent generation patterns, highlighting how linguistic factors and their corresponding cultural contexts systematically impact model outputs. Our benchmark and analysis provide a foundation for studying cross-lingual behavior in T2I generation and facilitate the development of more robust and inclusive models. Code and dataset are available at https://github.com/RISys-Lab/LingT2I.

Sources

Related papers