MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval

arXiv:2510.15543 · cs.CL, cs.AI, cs.IR, cs.MM · Submitted 2025-10-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval".

Jane: The gist: A modality composition awareness framework mitigates modality shortcut learning in unified multimodal large language models by enforcing structural relationships between composed and unimodal representations,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Now let's talk about who wrote this, or at least who is involved in the work behind "MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval." The authors are Wu, Cui, Hayakawa, Wang, Wakaki, and Mitsufuji.

Jane: It’s a team effort from researchers at Sony AI and other places who are looking at how to make these unified models more reliable. They’re focused on taking those large multimodal models that can handle many things at once and making sure they don't just get lazy about which information to trust.

Lu: These authors are clearly deep into the architecture, thinking about how the representations themselves should be structured, not just how the training data is fed in. They’re interested in creating a way for these models to understand the relationship between different types of inputs at a structural level.

Meng: I see they’re focused on solving a problem that arises naturally when you try to unify everything into one large architecture, which is pretty common right now. It sounds like they're trying to solve the inherent weakness of that unification.

Tom: Precisely, Meng. The core issue they are addressing is that while unified encoders are flexible, they can easily learn shortcuts by just leaning too heavily on one modality because the training loss gets minimized faster by exploiting that dominant signal.

Jane: So when we look at the title again, "Modality Composition Awareness," it tells us exactly what’s happening: they are making the model aware of how different modalities are composed together in a meaningful way.

Lu: It sounds like a very principled approach to regularization, moving away from just relying on standard contrastive learning objectives which don't explicitly care about the internal structure of the composition.

Tom: They introduce specific losses—the preference loss and the composition regularization objective—to enforce this structural awareness directly into the training process, rather than letting it emerge accidentally.

Meng: I wonder how complex those structural relationships are to model accurately, especially across different types of modalities like text and audio or image. It sounds like a lot of math to get right for practical use.

Jane: That’s the challenge, isn't it? They have to design losses that accurately reflect the desired behavior—that is, that the composition should be more distinct than any single part by enforcing that preference loss and alignment through composition regularization.

Tom: They are essentially teaching the model not just what inputs look like together, but *how* they should look together in a way that respects their individual components.

Lu: It’s an interesting direction because it suggests that for true multimodal reasoning, we need to teach the model about the interactions, not just present it with the final blended output.

The paper's summary: Tom: So let's get into what they actually propose in "MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval." They aren't just tweaking existing methods; they’re proposing a whole new training strategy to handle the modality shortcut problem.

Jane: The core of the paper is this framework where they introduce two complementary objectives. First, they have a preference loss that makes the combined multimodal embedding more discriminative than any of its unimodal counterparts.

Lu: That preference loss is formalized as L MCP, and it’s designed to actively discourage the model from relying on one modality by ensuring the composed similarity is higher than the individual unimodal similarities.

Meng: So they are quantifying that competition right there, making sure the model has to work harder to find a good representation when both modalities are present compared to when only one is. That’s a very concrete way to push it toward using both signals.

Tom: And then they have the second objective, the composition regularization objective, L MCR. This loss aligns those multimodal embeddings with prototypes that are built specifically from the unimodal parts of their inputs.

Jane: That second loss is designed to stabilize things by ensuring that even though it’s a composite embedding, it stays firmly anchored in the space defined by its individual components. It keeps everything tethered to what each piece contributes.

Lu: So, in short, they are combining making sure the combined result is superior to its parts with making sure that combined result respects the structure of those parts. It’s a two-pronged attack on modality shortcut learning.

Tom: That’s a pretty clear way to put it. They move beyond just training on the final retrieval task and build in these explicit structural constraints during training itself, which is what sets this work apart from just adding another layer of fine-tuning later.

Jane: It means they are making a strong argument that modeling the relationship between unimodal and multimodal representations structurally is necessary for robust performance, especially when facing distribution shifts.

Lu: The implication is that future retrieval systems won't just be about combining different encoders; they’ll need to explicitly model the interaction between those encoders.

Meng: For practical implementation, this means we have to be careful about tuning those alpha and beta parameters in their total objective function because if you set them wrong, you could end up fighting the system instead of helping it learn.

Tom: Right. And they show that when they use these combined losses, on out-of-distribution tasks they get noticeable gains, like plus five point nine percent on retrieval and five point four percent on zero-shot grounding compared to models trained only with contrastive learning.

Jane: So the summary is that MCA provides a principled way to mitigate modality shortcut learning by using both a preference loss and a composition regularization objective during training.

The paper's improvements: Tom: Now let's look at what they claim these methods actually achieve in terms of improvement over the previous approaches. They’re not just saying it’s better; they are showing quantifiable improvements on specific tasks.

Lu: The biggest takeaway is that under distribution shifts, MCA improves robustness, and this improvement holds up even when they test on new zero-shot tasks where the model has to generalize.

Jane: That robustness means the system doesn't completely break when it encounters a shift in the data distribution, which is a huge deal for real-world applications.

Meng: They demonstrated that even on in-domain benchmarks, all the different model variants trained with MCA end up hitting nearly identical accuracy levels. That shows they aren't just over-fitting to the training data.

Tom: That’s telling because it means these constraints are helping them learn a more generalized representation rather than just memorizing the specific examples they saw during training.

Lu: And when you look at their performance curves, the benefit of MCA doesn't show up immediately; you see it start after a few hundred steps, and that gap between out-of-distribution and in-domain performance widens within five hundred steps and continues to widen throughout the rest of the training.

Jane: So they are showing that this improvement builds over time, which makes sense because richer input information generally leads to stronger grounding, but MCA helps preserve that strength even when visual detail is degraded.

Tom: And they specifically highlight a situation where it really shines: in lower resolution settings, where the visual information is weaker and the risk of those shortcut learning issues is higher, MCA gives significant improvements by forcing the model to use those complementary textual cues.

Meng: It’s a practical application because we know that when visual data gets blurry or low-res, relying on text instructions becomes much more important for accurate understanding, and this paper validates that need.

Lu: The choice of how they implement the mixer doesn't seem to change things much for in-domain performance, but gated fusion actually gives them the best performance when it comes to those out-of-distribution settings.

Jane: So the overall improvement is that MCA provides a way to stabilize embeddings and keep them useful across different conditions, whether it’s in familiar data or completely new scenarios.

Conclusion: Tom: Alright, we've covered a lot about this paper "MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval." To wrap up, the main implication is that explicitly modeling the structural relationship between unimodal and multimodal inputs is a powerful tool against modality shortcut learning.

Jane: They’ve proven that by using both the preference loss to make combinations better than single parts and composition regularization to keep them anchored, you can get measurable gains in robustness under distribution shifts without losing performance on familiar data.

Lu: The big picture here is that for building next-generation multimodal retrieval systems, we need these explicit compositional constraints baked into the architecture from the start.

Meng: For a practical engineer, it suggests that when dealing with lower quality inputs, you need to be prepared to apply this kind of structural awareness to maintain accuracy.

Lalam: For me, as a model, I see this as an instruction set that helps me prioritize which modality is most important in any given situation so I don't default to the easiest path.

Tom: Exactly. It’s about making sure the AI doesn't just follow a shortcut because it's easier to do so during training on one type of data.

Jane: So we’ve seen how MCA works by showing that these two objectives work together effectively, leading to better performance on out-of-distribution tasks and more stable results overall.

Lu: Ultimately, this paper suggests that modeling the relationship between unimodal and multimodal inputs structurally is a key step toward building truly reliable AI systems.

Meng: I think the engineering focus should be on how to implement those compositional constraints so they are efficient enough for real-world use.

Lalam: And I think it’s helpful because it helps me learn to weigh the signals correctly, so I don't just default to one signal when things get ambiguous or messy.

Tom: That’s the gist of MCA, a solid piece of research that shows how structure can guide learning toward better generalization in multimodal retrieval.

Qiyu Wu, Shuyang Cui, Satoshi Hayakawa, Wei-Yao Wang, Hiromi Wakaki, Yuki Mitsufuji

Sony Group Corporation

cs.CL, cs.AI, cs.IR, cs.MM

Submitted: 2025-10-17

Updated: 2026-10-05

Importance score: 75/100

The gist: The gist: A modality composition awareness framework mitigates modality shortcut learning in unified multimodal large language models by enforcing structural relationships between composed and

Key concepts

Modality Shortcut Learning
This occurs when a unified model trained with contrastive learning ignores the complementary information from one modality (like text) and relies too heavily on another (like an image) during training. This makes the model brittle and unable to generalize well when faced with new data or distribution shifts.
Modality Composition Preference Loss ($L_{MCP}$)
This objective forces the embeddings of a combined multimodal input to be more distinct than any of its individual unimodal parts. By making the composed representation more discriminative, it actively discourages the model from taking shortcuts and encourages it to learn a true multimodal relationship.
Modality Composition Regularization ($L_{MCR}$)
This objective anchors the multimodal embedding to a space defined by its constituent unimodal embeddings. It ensures that when modalities are combined, the resulting representation remains closely aligned with the prototypes formed by each individual modality, stabilizing the composition.

Terminology

Summary

The gist: A modality composition awareness framework mitigates modality shortcut learning in unified multimodal large language models by enforcing structural relationships between composed and unimodal representations, leading to improved robustness under distribution shifts.

Introduction and Problem Statement

Multimodal retrieval aims to retrieve semantically relevant contents across multiple modalities such as text, image and audio, which is a fundamental task in various information fields The core ability of multimodal retrieval is to represent multimodal inputs in a shared and comparable embedding space A prevailing approach to this problem is to adopt unimodal encoders and align the encoded embeddings through contrastive learning (CL) However, with the rapid development of multimodal large language models (MLLMs), there has been growing interest in employing MLLMs as encoders for multimodal retrieval This flexibility also makes the model more prone to modality shortcut learning Since all modalities are jointly processed within a shared architecture, the training loss can be minimized by over-relying on the stronger modality signal, while ignoring the complementary one The toy example in Figure 1a illustrates that the model suffers from modality shortcuts when processing highly similar images, while ignoring the textual instruction The number of modalities increases, so does the risk of the model collapsing to rely on a single dominant modality, making explicit modality compositionaware constraints crucial for robust generalization

Proposed Modality Composition Awareness (MCA) Framework

The paper proposes a modality composition awareness (MCA) framework that models the structural relationship between multimodal and unimodal representations MCA consists of two complementary objectives from the preference and consistency perspective, respectively First, a preference loss enforces that the embeddings of a multimodal composition should be more discriminative than any of its unimodal counterparts, thereby discouraging modality shortcut learning Second, a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts The preference loss is formalized as L MCP = Ex,y+∼D [−X xm∈U(x) log e simθ(x,y+) e simθ(xm,y+)] This loss encourages the composed similarity higher than the unimodal similarities, ensuring that the model leverages complementary signals from multiple modalities rather than relying on one dominant modality The composition regularization objective is defined as L MCR = Ex,X∼D [− log e simmixθ,ϕ(x,U(x)) e simmixθ,ϕ(x,U(x'))] This loss enforces that the composed embedding remains anchored to the space formed by its constituent unimodal embeddings The overall training objective is formulated as L = L CL + α × LMCP + β × LMCR

Experimental Results and Analysis

Extensive experiments on both in-domain (IND) and OOD benchmarks demonstrate that the MCA improves robustness with OOD improvements under distribution shifts while maintaining IND performance The results show that when combined, the two losses yield substantially larger gains of +5.9% on OOD retrieval and 5.4% on zero-shot grounding, over the model trained with only contrastive learning On IND benchmarks, all model variants converge to almost identical accuracy levels The performance curves reveal that the benefits of MCA appear after a few hundred of steps, and the ODD gap emerges within 500 steps and continues to widen throughout training The richer input information yields stronger performance as expected, since richer visual detail provides more reliable visual grounding In lower resolution setting, where visual information is degraded and the risk of modality shortcuts is higher, MCA provides significant improvements by enforcing the use of complementary textual cues

Conclusion and Implications

In this work, we introduced modality composition awareness (MCA) for robust multimodal retrieval By explicitly modeling the relationship between unimodal and multimodal inputs, MCA incorporates two objectives: (1) modality composition preference, which discourages modality shortcuts by ensuring that composed representations are more discriminative than unimodal ones, and (2) modality composition regularization, which stabilizes embedding by aligning multimodal presentations with compositional prototypes Extensive empirical experiments across IND and OOD benchmarks demonstrate that MCA achieves consistent gains under distribution shifts and unseen zero-shot tasks, while maintaining comparable IND accuracy These results highlight MCA as an effective principle for mitigating modality shortcut problem when using MLLMs as unified encoder and improving generalization in multimodal retrieval The richness of the input information dictates the weighting required for MCA to be effective, suggesting that stronger regularization is needed for weaker inputs to preserve robust modality composition The choice of mixer implementation has little effect on IND benchmarks, but gated fusion achieves the best performance in OOD settings In summary, MCA effectively reduces the modality shortcut in diverse cases from multiple benchmarks

--- Page 1 ---

Manuscript

MCA: MODALITY COMPOSITION AWARENESS FOR ROBUST COMPOSED MULTIMODAL RETRIEVAL Qiyu Wu1 Shuyang Cui1 Satoshi Hayakawa1 Wei-Yao Wang1 Hiromi Wakaki1 Yuki Mitsufuji1,2

--- Page 2 ---

Manuscript

Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production Despite the success of separate-encoder approaches like CLIP align modalityspecific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts

--- Page 3 ---

Manuscript

Multimodal retrieval, which aims to retrieve semantically relevant contents across multiple modalities such as text, image and audio, is a fundamental task in various information fields Its applications span a wide range of domains, such as text-vision retrieval(Huynh et al., 2025; Wang et al., 2021), music retrieval(Doh et al., 2023), product search(Goenka et al., 2022; Zhu et al., 2024b), and multimodal retrieval-augmented generation (Yasunaga et al., 2023; Ghosh et al., 2024; Yang et al., 17> The core ability of multimodal retrieval is to represent multimodal inputs in a shared and comparable embedding space A prevailing approach to this problem is to adopt unimodal encoders and align the encoded embeddings through contrastive learning (CL) Models following this separate-encoder paradigm, such as CLIP (Radford et al., 2021) and CLAP (Elizalde et al., 2023), have demonstrated the effectiveness of CL, achieving strong performance across various multimodal retrieval tasks On the other hand, with the rapid development of multimodal large language models (MLLMs) (Alayrac et al., 2022; Li et al., 2023; Bai et al., 2023; Chu et al., 17> there has been growing interest in employing MLLMs as encoders for multimodal retrieval (Jiang et al., 17; Zhang et al., 17; Huang et al., 17; Jiang, 18> Unlike separate-encoder frameworks, MLLMs are capable of processing inputs from different modalities, as well as their compositions, within a unified architecture However, this flexibility also makes the model more prone to modality shortcut learning Since all modalities are jointly processed within a shared architecture, the training loss can be minimized by over-relying on the stronger modality signal, while ignoring the complementary one A toy example in Figure 1a illustrates that the model suffers from modality shortcuts when processing highly similar images, while ignoring the textual instruction

--- Page 4 ---

Manuscript

Figure 1: Illustration of modality shortcut problem in multimodal retrieval. (a) is a conceptual example where the red-highlighted text descriptions are ignored, leading the model to follow a shortcut path, instead of the expected composed representation (b) and (c) are t-SNE visualizations of composed queries, unimodal queries and targets, that randomly sampled from CIRR dataset, under vanilla CL and our MCA, respectively tSNE implementation details are introduced in §A.6 standing on flat asphalt”. Illustration of learned representation with vanilla CL in Figure 1b also demonstrates that composed queries are poorly separated and often cluster near text-only queries, indicating that the model relies on modality shortcut instead of learning robust composed representations Hence, the architecture shift from separate-encoder to unified-encoder makes the direct application of conventional CL objective to unified MLLM encoders limiting to robustness Furthermore, the modality shortcut problem naturally extends to scenarios involving multiple modalities, particularly as recent work aims to unify various modalities into a single framework (Girdhar et al., 2023; Zhu et al., 17; Xu et al.

Improvements for AI systems

  1. Bold header: Robust Composed Retrieval via MCA

This improvement enables MLLM-based retrievers to maintain performance under distribution shifts by explicitly modeling structural relationships between multimodal and unimodal representations, mitigating modality shortcut learning.

  1. Bold header: Enhanced Out-of-Distribution Generalization

The system can achieve significant improvements on OOD benchmarks, specifically in retrieval and zero-shot grounding tasks, demonstrating robust modality composition awareness when facing unseen compositions or domains.

  1. Bold header: Improved Performance Under Input Degradation

MCA allows the model to maintain performance even when input information is degraded, as shown by results where lower resolution setting, where visual information is degraded and the risk of modality shortcuts is higher, MCA provides significant improvements.

Sources

Related papers