Evaluating Perspectival Biases in Cross-Modal Retrieval
cs.IR, cs.CL
Submitted: 2025-10-30
Updated: 2026-08-30
Project page: https://commoncrawl.github.io/cc-crawl-statistics/plots/languages
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query.
Terminology
Abstract
Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: deviations shaped by linguistic prevalence and cultural associations. We introduce the Cross-Cultural, Cross-Modal, Cross-lingual Multimodal (3XCM) benchmark to isolate these effects. Results from our studies indicate that, for image-to-text retrieval, models tend to favor entries from prevalent languages over those that are semantically faithful. For text-to-image retrieval, we observe a consistent "tugging effect" in the joint embedding space between semantic alignment and language-conditioned cultural association. When semantic representations are insufficiently resolved, particularly in low-resource languages, similarity is increasingly governed by culturally familiar visual patterns, leading to systematic association bias in retrieval. Our findings suggest that achieving equitable multimodal retrieval necessitates targeted strategies that explicitly decouple language from culture, rather than relying solely on broader data exposure. This work highlights the need to treat linguistic and cultural biases as distinct, measurable challenges in multimodal representation learning.
Sources
- Learning Transferable Visual Models From Natural Language Supervision
- Qwen2.5-VL Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions
- Mitigating Language Bias in Cross-Lingual Job Retrieval: A Recruitment Platform Perspective
- MIRACL-VISION: A Large, multilingual, visual document retrieval benchmark
- Fairness and Bias in Multimodal AI: A Survey
- RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
- GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG