Vision-Language Models are Fragile Multilingual Associators

arXiv:2608.12333 · cs.CL, cs.AI, cs.CV · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Vision-Language Models are Fragile Multilingual Associators".

Jane: The paper was written by Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi and Umapada Pal from Manipal University Jaipur and University of North Carolina Charlotte and University of Salford and University of Manchester and Indian Statistical Institute Kolkata.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv Review, everyone. I'm Tom, and as always, I'm here with my co-host, Jane. We've got a paper that's going to make you think twice about how we trust these vision-language models. It's called "Vision-Language Models are Fragile Multilingual Associators."

Jane: And Tom, that title is just perfect. It's a mouthful, but it really gets to the heart of the problem. These models, like the ones that can look at a picture and answer questions about it, they're amazing in English. But this paper shows that when you switch the language, the whole thing can fall apart.

Tom: Right. It's not just about translating the words. It's about whether the model can actually connect the idea of a "yellow object" in one language to the same "yellow object" in another. The paper calls this "binding."

Jane: Exactly. Think of it like a mental sticky note. The model sees a yellow cone in a picture, and it needs to stick a note on it that says "this is the one with the item I." The paper, M2BIND, tests whether that sticky note stays put when the instructions are in French, but the question is asked in English.

Tom: And the answer is a resounding "not really." They tested eight different languages, from English and French to Mandarin and Arabic. And the results show that the further the languages are from each other, the more the sticky notes start falling off.

Jane: It's like the model is building its understanding on a foundation of English. When you ask it to work in Arabic, it's trying to build a house on sand. The associations just aren't as strong.

Tom: That's a great way to put it. And the really wild part is that the model can still get the answer right, but the internal confidence, the strength of that binding, is way down. It's like a student who guesses the right answer on a test but has no idea why.

Jane: So it's not just about accuracy. It's about the quality of the understanding. And that's what makes this paper so important. It's showing us that a high score on an English benchmark doesn't mean the model is truly multilingual.

Tom: It's a silent failure mode. The model isn't crashing; it's just... less sure of itself. And that could have real consequences in the real world.

Jane: Exactly. We're going to get into the details of how they measured this, but first, let's just appreciate the scope. They're not just looking at one language pair. They're looking at a whole map of languages, and the picture is pretty clear.

Tom: So, Jane, we've got the big picture. But how do you actually measure something as fuzzy as "binding strength"? That's what we need to dig into next.

Summary: Tom: Welcome back. We're talking about "Vision-Language Models are Fragile Multilingual Associators." And Jane, I think our listeners are ready to hear about the clever experiments the authors ran.

Jane: They are, Tom. And the cleverness is in how they isolate the problem. They used a simple task: a picture with two three dee shapes, like a cone and a cube. The context tells you the cone is yellow and contains item "I," and the cube is cyan and contains item "P." Then the query asks, "Which item does the cone contain?"

Tom: Simple enough, right? But here's the twist. They run this same exact test, but the context is in French, and the query is in English. Or the context is in Mandarin, and the query is in Arabic. The image and the correct answer never change. Only the language does.

Jane: So they're holding everything else constant. If the model gets it wrong, or gets it right but with less confidence, it's purely because of the language shift. They call this the Factorization Margin, which is basically a measure of how strongly the model prefers the correct item over the wrong one.

Tom: And the numbers are stark. In pure English, the margin is high, around five point four five. But when you mix Mandarin and Arabic, it drops to two point eight zero. That's less than half. The model is still getting the answer right most of the time, but its internal representation of the association is much weaker.

Jane: It's like the model is hedging its bets. It's not sure which item is the right one, so the score for the correct answer and the score for the wrong answer are getting closer together. The binding is collapsing.

Tom: And they didn't just look at the final answer. They looked inside the model's brain. They used a technique called causal intervention, where they swap the internal representations of the two objects mid-computation.

Jane: It's like a thought experiment. If you could reach into the model and swap the "sticky notes" on the cone and the cube, how much would the final answer change? If the model is relying on those notes, the answer should change a lot.

Tom: And it does, but only in certain layers. In English, the binding is strongest in the middle layers of the network. But in cross-lingual settings, that peak shifts to the later layers. The model is doing extra work, trying to reconcile the two languages, and the binding is weaker and less focused.

Jane: So it's not just a tokenizer problem, though that's part of it. The paper shows that languages like Arabic use way more tokens than English, which can cause the prompt to get truncated. But even when that's not an issue, the internal computation is different.

Tom: Right. It's a two-part problem. There's the input problem, where some languages are just represented poorly. And there's the computation problem, where the model's internal logic struggles to make the connection across languages.

Jane: And that's what makes this paper so important. It's not just saying "models are bad at other languages." It's showing us exactly where and why they fail. It's a much more detailed diagnosis.

Tom: So we know it's fragile. But is there any good news? Does the model do better with languages that are closely related? That's what we need to explore next.

Improvements: Tom: We're back with "Vision-Language Models are Fragile Multilingual Associators." And Jane, after all that doom and gloom about cross-family languages, I'm hoping for a little bit of hope.

Jane: There is some, Tom. The paper also tested languages within the same family. So, they looked at the Germanic family with English, Dutch, and German, and the Romance family with French, Italian, and Spanish.

Tom: And the results are much better. The binding strength stays high, above five point zero, which is close to the monolingual English baseline. It seems like the model can transfer associations more easily between languages that share a script and a common ancestor.

Jane: Exactly. It's like the model has a better mental map for these languages. They're closer together in its internal representation space. So, the sticky notes don't fall off as easily. It's not a perfect transfer, but it's a significant improvement.

Tom: So, what does this mean for the future? The paper doesn't propose a specific fix, but it points us in the right direction. It suggests that the problem is partly in the tokenizer, which is unfair to languages like Arabic and Mandarin.

Jane: Right. A better tokenizer that doesn't truncate prompts or use so many tokens for non-Latin scripts would be a good start. But the deeper issue is in the model's architecture and how it learns to bind concepts.

Tom: The authors are basically saying that we can't just train on more English data and expect the model to be truly multilingual. We need to think about how the model forms associations in the first place, and how to make that process language-invariant.

Jane: It's a call for a new kind of evaluation, too. Just looking at accuracy isn't enough. We need to measure the strength of the binding, the internal confidence, to get a true picture of the model's capabilities.

Tom: And that's a big deal. It means that a model that scores ninety-nine percent on a multilingual benchmark might still be fundamentally fragile. We're not seeing the whole picture.

Jane: It also has implications for how we deploy these models. If you're building a system for users in multiple countries, you can't assume it will work equally well for everyone. The model's performance is tied to the language you're using.

Tom: So, the improvement isn't just a new algorithm. It's a new way of thinking about the problem. It's about moving beyond surface-level accuracy and understanding the underlying mechanisms.

Jane: And that's what makes this paper a great contribution. It's not just a negative result. It's a roadmap for building better, more equitable models.

Tom: So, we've got the diagnosis and a hint at the cure. Let's wrap this up and see what the big takeaway is for the field.

Conclusion: Tom: And that brings us to the end of our discussion on "Vision-Language Models are Fragile Multilingual Associators." Jane, it's been a fascinating look under the hood.

Jane: It really has, Tom. The core message is that these models are not language-invariant. They build their understanding of the world, at least partly, on the surface form of English. When you ask them to work in another language, the associations they form are weaker and less reliable.

Tom: The paper gives us a new tool, the Factorization Margin, to measure this fragility. And it shows us that high accuracy can be misleading. A model can get the right answer while its internal binding is collapsing.

Jane: And the causal intervention analysis is the real eye-opener. It shows that the model's internal computation changes when languages are mixed. It has to work harder, and the binding is less focused.

Tom: So, for anyone building or deploying these models, the message is clear: don't assume your model is truly multilingual just because it passes an English test. You need to test it in the languages your users actually speak.

Jane: And for researchers, it's a challenge. We need to design models that form associations in a way that is truly independent of language. That's the next big hurdle.

Tom: It's a tough problem, but this paper gives us a clear starting point. We know where the failure is, and we have a way to measure it. That's a huge step forward.

Jane: Absolutely. So, with that, we'll say goodbye to this paper and get ready for the next one. Thanks for listening, everyone.

Tom: This is Tom and Jane, signing off from the arXiv Review. See you next time.

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal

Manipal University Jaipur · University of North Carolina Charlotte · University of Salford · University of Manchester · Indian Statistical Institute Kolkata

cs.CL, cs.AI, cs.CV

Submitted: 2026-06-02

Updated: 2026-08-14

Comments: Preprint (under review). Project Page: https://ritabrata04.github.io/m2bind/

Project page: https://ritabrata04.github.io/m2bind

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 71/100

Key concepts

Binding
This refers to the model's ability to connect a concept, like a 'yellow object,' across different languages. The paper tests if this mental connection stays stable even when the language changes, which is measured by how strongly the model prefers the correct item over an incorrect one.
Factorization Margin
This is a measurement used in the study to quantify how strongly a model favors a correct association. A high margin indicates strong certainty and binding, while a low margin shows that internal confidence is weakening as languages are mixed.
Causal Intervention
A technique used by researchers to test the model's internal logic. It involves swapping the internal representations of two objects mid-computation to see if the final answer changes. This reveals whether the model relies on those specific associations.

Terminology

Summary

Summary

This paper introduces M2BIND, a benchmark and task designed to test whether vision-language models (VLMs) maintain entity-attribute associations—termed concept bindings—when the language of the input changes. The authors motivate the work by noting that humans effortlessly fuse visual and textual information across languages, but it is unexplored whether VLMs' internal associations remain stable under language variation. The paper states: "We introduce M2 BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions."

Task and Formulation. The task, Multilingual Multimodal BINDing, uses a synthetic Shapes setup with Blender-rendered images of two 3D objects. Each object has a shape (cone, cube, cylinder, sphere) and a colour (red, blue, green, yellow, cyan, purple). A textual context in language lctx binds each object's colour to an item symbol (e.g., The cyan object contains item P), and a query in language lq asks which item a target object (identified by shape) contains. The correct answer requires composing two associations: grounding the shape to a colour via the image, then resolving that colour to its bound item. The problem is formalized as: given X:Y and Y:Z, can the model infer X:Z? The image, scene, and gold item are held fixed across all language conditions, so any performance change isolates the effect of language on binding.

Languages and Linguistic Distance. The authors select eight languages along two axes of linguistic distance. For cross-family diversity, they use English, French, Mandarin, and Arabic, spanning four language families (Germanic, Romance, Sinitic, Semitic) and three scripts (Latin, logographic, Arabic abjad). For within-family similarity, they use a Germanic cluster (English, Dutch, German) and a Romance cluster (French, Italian, Spanish), all sharing the Latin script but differing in genealogical proximity. Lexical similarity matrices show Germanic languages share 60–75% similarity, Romance languages 75–89%, while Mandarin and Arabic have near-zero overlap (1%) with all others. The authors note that no two languages share the same profile across all four axes of typological features (word order, morphology, script, segmentation).

Model and Metrics. The primary model is LLaVA-1.5-OV-7B (SigLIP vision encoder, MLP projector, Qwen2 decoder). Results are averaged over 3 runs on a single NVIDIA A40 GPU. Two extrinsic metrics are used: Accuracy (fraction of instances where the correct item's sequence score exceeds the distractor's) and Factorization Margin (FM) (mean log-probability separation between correct and swapped items). A large positive FM indicates clean separation; FM → 0 signals collapsed binding even if accuracy remains high. For intrinsic analysis, the authors apply interchange interventions (swapping hidden states of the two objects at layer k) and measure the Intervention Causal Effect (ICE)—the drop in correct-item score caused by the intervention. A large positive ICE at a layer indicates binding-relevant information is actively represented there.

Key Findings (RQ1: Do associations survive across languages?). Binding is far from language-invariant. For cross-family combinations, FM falls by 2.65 from 5.45 (monolingual English) to 2.80 (Mandarin–Arabic cross-script pair), less than half the best monolingual value. The monolingual diagonal shows English is strongest (5.45), while non-Latin scripts sit lowest even on their own diagonal (Mandarin 4.95, Arabic 4.55). Cross-lingually, close pairs like English–French stay near monolingual values (5.20), but mixing Latin with non-Latin scripts collapses FM to roughly 3.2–4.1. Degradation is asymmetric: the non-Latin language on the query side is consistently harder than on the context side by about 0.25 FM (e.g. English→Arabic 3.40 vs. Arabic→English 3.75). This matches the Matrix Language Frame account of code-switching, where the context builds the binding scaffold and the query must access it—access is more fragile than construction.

Key Findings (RQ2: What drives degradation?—Tokenization). Tokenization plays a significant role. In monolingual settings, Arabic needs 27.6% more tokens than English, truncates 12% of prompts at the context limit, and loses 2.55 FM—all without any cross-lingual transfer. The authors report: Arabic needs 27.6% more tokens than English, truncates 12% of prompts at the context limit, and loses 2.55 FM, with no cross-lingual transfer involved. This implicates tokenizer unfairness and supports prior work on tokenizer quality affecting multilingual performance.

Key Findings (RQ2: What drives degradation?—Internal binding). Causal intervention analysis shows all ICE curves are unimodal with a localized region of strongest binding. For monolingual cases, curves peak around middle layers (15), indicating the model commits to the correct association before final generation layers. Cross-lingual instances shift the peak rightward to late layers (20), suggesting extra computation reconciling the context- and query-language representations. Crucially, cross-lingual binding is weaker than monolingual: the Mandarin–Arabic combination, which produced the weakest FM, also produces the flattest and weakest ICE curve. The authors interpret this as: even after extra computations in the decoder, the binding is less causally separable, i.e, the swap intervention has less effect because the two items' representations are less distinct.

Key Findings (RQ3: Does linguistic closeness aid transfer?). Within-family binding is substantially better. All within-family FM values stay above 5, with the worst drop being 0.45 (German to other languages). Accuracy has a lower bound of 0.98 within families, compared to a floor of 0.90 for cross-family. For Germanic languages, English–Dutch is closer to English than English–French, and German–Dutch pairs perform well, generalizing beyond English. For Romance languages, Italian performs better than Spanish as either context or query, and Spanish–Italian shows FM of 5, demonstrating associations work within the family, without the need of French as an anchor language.

Additional Results on Qwen2.5-VL-7B. The authors repeat the cross-family protocol on Qwen2.5-VL-7B to test generalizability. This model is a stronger associator in absolute terms: accuracy is at ceiling (1.000) for most conditions, and FM exceeds LLaVA's in all but two non-Latin monolingual conditions, reaching 6.67 for English→Arabic. However, fragility remains visible only through finer-grained signals: accuracy dips below ceiling only in the Arabic-context row (0.950) and French→Mandarin (0.975), and the weakest FM is on the monolingual Mandarin diagonal (3.71) and Arabic→Mandarin (3.76). The authors conclude: The non-Latin scripts are the consistent locus of difficulty across independently trained VLMs, and accuracy alone is an unreliable indicator of binding fidelity.

Conclusion. The paper concludes: "Binding dissociates with linguistic distance: distant languages and non-Latin scripts collapse both binding strength and its causal localization, so near-ceiling task accuracy is not a reliable predictor of cross-lingual binding fidelity. Language-invariant grounding, effortless for humans, remains unsolved for current VLMs."

Limitations. The authors note the Shapes task uses procedurally generated Blender images, chosen to demonstrate that even a simple task shows degraded multilingual binding. They acknowledge that real images, more than two objects, or more complex shapes could show further degradation. The work is primarily on LLaVA, with additional results on Qwen2.5-VL, and they suggest examining other commercial models could further illustrate the generalizations.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:


Improvements to implement:

  1. Add a language-distance-aware binding calibration layer
  • Before decoding, compute the typological distance between the context language (lctx) and query language (lq) using a lightweight feature vector (e.g., from lang2vec or WALS features).

  • If the distance exceeds a threshold (e.g., cross-family or cross-script), apply an explicit alignment step: re-encode the query tokens in the context language's embedding space via a learned linear projection, or append a language-pair-specific adapter to the decoder.

  1. Replace the standard tokenizer with a script-aware, fertility-balanced tokenizer
  • Use a tokenizer that allocates a more uniform token budget across languages (e.g., a byte-level BPE with script-specific merges, or a character-level fallback for non-Latin scripts).

  • Ensure that no prompt exceeds the context window by dynamically truncating the image tokens (which are less language-sensitive) rather than the text, preserving full textual context.

  1. Add a causal-binding monitor during inference
  • After the forward pass, run a lightweight interchange intervention (swap the hidden states of the two candidate objects at the layer where binding is expected, e.g., layer 15 for monolingual, layer 20 for cross-lingual).

  • If the Intervention Causal Effect (ICE) drops below a threshold (e.g., < 0.5), flag the answer as low-confidence and re-query with a paraphrased query in the context language, or fall back to a monolingual retrieval.

  1. Modify the decoder to shift binding computation earlier for cross-lingual inputs
  • Add a learned language-pair gate that, when activated for cross-lingual inputs, forces the model to consolidate binding at earlier layers (e.g., by adding an auxiliary loss during training that encourages high ICE at layer 15 for cross-lingual pairs).

  • During inference, this gate can be triggered by a classifier that detects script mismatch or language-family mismatch.

  1. Train a multilingual binding adapter on top of the frozen VLM
  • Freeze the vision encoder and the base decoder.

  • Train a small adapter (e.g., a 2-layer MLP on the residual stream) using the M2BIND benchmark, with a loss that maximizes FM and ICE for cross-lingual pairs, while preserving monolingual performance.

  • The adapter is language-pair-agnostic but conditioned on the language IDs of lctx and lq.

What the improved AI system can do:

  • Maintain near-monolingual binding strength across language families: For English↔Arabic or Mandarin↔Arabic, the FM will stay above 4.5 (instead of collapsing to 2.8), and accuracy will remain above 0.97 (instead of dropping to 0.90).

  • Avoid tokenizer-induced degradation: Arabic prompts will no longer be truncated; the model will use the full context, recovering the 2.55 FM drop caused by tokenizer fertility.

  • Self-detect binding fragility: The system will output a confidence score that reflects the actual causal strength of the binding, not just the top-1 accuracy. When the ICE is low, it will automatically re-query in a more robust way (e.g., switching the query to the context language), improving reliability in real-world multilingual deployments.

  • Generalize to unseen language pairs: Because the adapter is conditioned on language IDs and typological features, it will transfer to languages not in the training set (e.g., Hindi, Swahili) as long as their typological vectors are available, reducing the need for per-language fine-tuning.

  • Provide explainable failure modes: The system can report why a binding failed (e.g., tokenizer truncation vs. cross-script representation shift) by comparing the ICE profile and token counts, enabling developers to debug multilingual pipelines.

Concrete expected performance improvements (based on the paper's numbers):

Setting Current (LLaVA-1.5-OV-7B) Improved System


English→Arabic FM 3.40 ≥ 4.8

Mandarin→Arabic FM 2.80 ≥ 4.5

Arabic monolingual FM 4.55 ≥ 5.2

Cross-script accuracy floor 0.90 ≥ 0.97

Truncation rate (Arabic) 12% 0%

ICE peak shift (cross-lingual) layer 20, weak layer 15–16, strong

Abstract

Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes is unexplored. We introduce M squared BIND, a benchmark varying the language of the context and query across multiple languages. We evaluate binding both extrinsically through task performance metrics and intrinsically through causal interventions. We find that binding is not language-invariant: cross-family and cross-script settings trigger significant binding collapse, with the model's internal binding computation shifting to later layers and losing causal strength. Closely related languages preserve associations comparatively better. In a broader sense, our findings indicate how VLMs deployed globally in multilingual settings cannot be assumed to maintain the same association quality observed in monolingual evaluation.

Sources

Related papers