Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Neither Here Nor There".
Jane: Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the authors of "Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders." It includes Divyansh Pathak, Prashant Kodali, and Jasabanta Patro from the Indian Institute of Science Education and Research.
Jane: That's a great team with deep expertise in multilingual modeling; it tells us this research has a strong foundation in how these massive models are structured.
Lu: Their work is building on prior studies, looking at how English and Hindi representations align well separately, but then seeing where the code-mixed signals actually connect within that shared space.
Meng: What I find interesting is that they are specifically focusing on the representational dynamics rather than just surface-level performance metrics which is a deeper dive into model behavior.
Lalam: It means they aren't just telling us *what* the model gets right, but *how* it’s thinking about that mixed text at a structural level.
The paper's summary: Tom: To summarize what the paper covers, they first constructed a unified trilingual corpus of parallel English, Hindi in Devanagari script, and Hindi–English code-mixed sentences. This gave them the material to test their hypotheses.
Jane: And the core finding they present is that even when standard models align English and Hindi representations effectively on their own, the code-mixed inputs still show a weak connection to either of those constituent languages internally.
Lu: The key observation here is that code-mixed representations remain loosely anchored to both languages under standard pretraining conditions, which leaves the relationship between the mixed input and its monolingual parts underexplored.
Meng: So, it's not a simple failure of translation; it's a structural issue in how the model maps these different linguistic forms into its embedding space when they are mixed together.
Lalam: This means that simply feeding the model more code-mixed data won't automatically fix this fundamental disconnect; we need to address the representation geometry itself.
The paper's improvements: Tom: The paper suggests a way forward by looking at how continued pretraining on code-mixed data specifically impacts alignment, and they found a trade-off. They noted that while continued pretraining on code-mixed data improves English–code-mixed alignment, it comes at the cost of English–Hindi alignment.
Jane: That trade-off is really telling; it shows that adaptation doesn't always improve everything equally, and it can actually weaken a different language pair in this specific context.
Lu: They also pointed out that the script choice matters quite a bit; when Hindi is Romanized during pretraining, the alignment between Hindi and code-mixed text actually degrades.
Meng: From an engineering standpoint, knowing that script preservation helps alignment suggests we need to be careful about how we prepare our input data for models trained on mixed scripts.
Lalam: This points toward a solution where we can guide the model's attention more precisely, maybe through a specific post-training objective that balances these conflicting alignment signals between English and Hindi.
Conclusion: Tom: So, to wrap up this look at "Neither Here Nor There," the authors conclude that while code-mixed pretraining helps English–code-mixed alignment, it's not a complete solution on its own. They propose an explicit trilingual post-training alignment stage using an objective designed to pull semantically equivalent sentences closer together in the shared embedding space.
Jane: That leads us to their main contribution: introducing the Cross-Lingual Alignment Score, or CLAS, as a metric to measure that balanced cross-lingual alignment across all language pairs.
Lu: This stage is designed to encourage more balanced cross-lingual alignment between English, Hindi, and code-mixed representations while trying to keep monolingual consistency strong.
Meng: For practical deployment, this suggests we can use this explicit alignment objective during the fine-tuning phase to ensure that when we deploy a model for a specific application, like sentiment analysis on code-mixed text, its internal logic is robust across all three languages involved.
Lalam: I think the biggest implication for culture is that these improved representations mean our AI can finally handle the rich, nuanced ways people communicate in real-world social media settings without losing meaning or accuracy in any single language component.
Indian Institute of Science Education and Research, Bhopal · Microsoft Corporation
cs.CL
Submitted: 2026-03-20
Updated: 2026-10-01
Comments: Accepted EMNLP Findings 2026
Code: https://github.com/devanshg27/cm_translation
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 78/100
The gist: Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally — or whether those
Key concepts
- Code-Mixed Representations
- These are the internal mathematical ways multilingual models store text that mixes two or more languages. The researchers found these representations are often weakly tied to both languages initially, primarily leaning on the English representation space.
- EN–CM Alignment Trade-off
- When models are pretrained on mixed language data, improving alignment between English and code-mixed text comes at the expense of alignment between English and Hindi. This suggests that code-mixed data signals are strongest for the EN–CM pair.
- Trilingual Post-Training Alignment
- This is a proposed method where models are further trained using a new objective to push embeddings of sentences from English, Hindi, and code-mixed text closer together in the shared space. This aims to create more balanced alignment across all three language pairs.
Terminology
Summary
Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally — or whether those representations meaningfully connect to the constituent languages being mixed.
The gist: Code-mixed adaptation of multilingual encoders improves EN–CM alignment at the cost of EN–HI, reorienting the anchor choice for cross-lingual alignment and leading to better cross-lingual performance on downstream tasks.
Dataset Construction
The study constructed a unified trilingual corpus by merging three publicly available datasets: CM-En Parallel (Dhar et al., 2018), PHINC (Srivastava and Singh, 2020), and the LinCE 2021 multiview Hinglish dataset (Aguilar et al., 2020). This corpus comprises parallel English, Hindi (Devanagari script), and Hindi–English code-mixed sentences. The final dataset contains 21,139 aligned sentence triples. Since the first two datasets did not provide Hindi translations, missing sentences were generated using the Google Translate API and verified using the IndicTrans2 model to ensure consistency.
Cross-Lingual Representation Analysis
The researchers examined how multilingual encoders represent code-mixed text relative to their monolingual counterparts using interpretability tools such as Centered Kernel Alignment (CKA), token-level saliency, and entropy-based uncertainty analysis. Key observations revealed that code-mixed representations remain weakly anchored to both languages
under standard pretraining. Specifically, code-mixed representations are largely explained by the English representation space,
while native-script Hindi provides complementary signals that reduce representational uncertainty.
Impact of Code-Mixed Adaptation
Continued pretraining on code-mixed data was investigated to understand its effects on alignment. The findings indicated a systematic trade-off: continued pretraining on code-mixed data improves English–code-mixed alignment at the cost of English–Hindi alignment.
After adaptation, the "alignment ordering inverts to: EN ↔ CM > EN ↔ HI > HI ↔ CM, suggesting that codemixed data provides alignment signals primarily between English and code-mixed settings. Furthermore, script choice was found to be significant:
When Hindi is Romanized during pretraining, HI↔CM alignment degrades, whereas
preserving Hindi in Devanagari improves alignment."
Proposed Trilingual Post-Training Alignment Stage
Motivated by these dynamics, the authors introduced a trilingual post-training alignment stage. This stage adapts pretrained multilingual encoders using the objective:
L = Lbase + λ · Lalign (3)
where Lalign encourages semantically equivalent sentences across EN, HI, and CM to occupy nearby regions in the shared embedding space.
The Cross-Lingual Alignment Score (CLAS) was proposed as a metric to measure balanced cross-lingual alignment across language pairs. Experiments showed that incorporating this trilingual supervision leads to more balanced cross-lingual alignment across language pairs while largely preserving strong monolingual consistency.
Downstream Task Validation
The effectiveness of the proposed stage was evaluated on two code-mixed classification tasks: sentiment analysis and hate speech detection. The results demonstrated that trilingual alignment encourages stable multilingual representations and improves cross-lingual robustness,
leading to improved cross-lingual prediction consistency across CM, EN, and HI inputs for both tasks. For instance, in sentiment analysis, the trilingual alignment applied to Hing-mBERT improved the consistency score from 0.3283 to 0.4464 when training on code-mixed data. Similarly, for hate speech detection with Hing-RoBERTa, it improved the consistency score from 0.6124 to 0.6720 when training on Hindi, suggesting that trilingual alignment is more effective when the finetuning language is Hindi.
Conclusion
The study concludes that while code-mixed pretraining improves EN-CM alignment, the explicit trilingual alignment objective provides consistent and additive benefits over code-mixed pretraining alone,
particularly in achieving a more balanced cross-lingual generalization across English, Hindi, and code-mixed representations. This approach not only yields better cross-lingual performance on downstream tasks but also mitigates the degradation of EN–HI alignment.
Limitations
The study notes that the analysis focuses on Hindi–English code-mixing as a representative case study, though the mechanisms are expected to generalize. Future work is suggested to extend the analysis to additional languages and scale the alignment stage with larger and more diverse code-mixed corpora. Additionally, limitations include reliance on automatic translations for corpus construction and focusing primarily on encoder-based models rather than decoder-based LLMs.
Ethical Considerations
The authors noted that the study involves social media text potentially containing offensive or harmful language, which was included solely for research purposes to study linguistic phenomena in code-mixed settings.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems, based on the findings of this research, and what those improved systems could achieve:
)
-
Improvements in Multilingual Encoder Pre-training Objectives:
-
Incorporation of a Trilingual Post-Training Alignment Stage (Trilingual Post-Training Alignment Objective):
-
Enhanced Cross-Lingual Representation Consistency Scoring (Cross-Lingual Alignment Score - CLAS):
-
Improved Robustness and Stability in Code-Mixed Classification Tasks:
-
The AI system will be trained using a modified pre-training objective that explicitly incorporates parallel sentence supervision across English, Hindi (Devanagari), and Romanized code-mixed text. The objective function will be:
L = Lbase + λ · Lalign
- This improved system can achieve:
2.1. More balanced cross-lingual alignment across all language pairs (EN–HI, EN–CM, HI–CM).
2.2. Mitigation of the degradation in the English–Hindi alignment that often occurs when adapting models to code-mixed data (a key trade-off identified in Section 4).
-
The system will utilize a new metric, the Cross-Lingual Alignment Score (CLAS), to evaluate its internal representation geometry, moving beyond simple task performance metrics.
-
The AI system can achieve:
3.1. A quantifiable measure of how balanced
the model's shared embedding space is across its constituent languages.
3.2. Identification of which specific training regimes (e.g., pre-training on CM data vs. post-training alignment) yield the most robust, bidirectional cross-lingual performance, guiding optimal fine-tuning strategies for multilingual models.
-
The system will leverage interpretability techniques (CKA and Entropy-Based Uncertainty Reduction) during development to understand internal representation dynamics.
-
This capability allows researchers to:
5.1. Determine the semantic subspace
through which code-mixed text is processed, confirming if the model relies excessively on an English-dominant space or if native script signals (like Devanagari Hindi) are effectively providing complementary, uncertainty-reducing information.
-
The system will demonstrate improved cross-lingual consistency on downstream tasks (Sentiment Analysis and Hate Speech Detection).
-
Specifically, the AI system will exhibit:
7.1. Higher Macro-F1 scores and superior performance stability when processing code-mixed inputs compared to standard models, especially in the more complex Hindi–English setting.
7.2. Enhanced robustness in hate speech detection by ensuring that classification consistency is maintained across different language training conditions (CM, EN, HI), leading to more reliable safety evaluations.
Abstract
Multilingual encoder-based language models are widely used for code-mixed analysis, yet their internal representations of code-mixed inputs -- and their relationship to the constituent languages -- remain poorly understood. Using Hindi-English as a case study, we construct a unified trilingual corpus of parallel English, Hindi (Devanagari), and Romanized code-mixed sentences. We then probe cross-lingual representation alignment in standard multilingual encoders and their code-mix-adapted variants using CKA, token-level saliency, and entropy-based uncertainty analysis. We find that while standard models align English and Hindi well, code-mixed inputs remain loosely connected to either language -- and that continued pre-training on code-mixed data improves English-code-mixed alignment at the cost of English-Hindi alignment. Interpretability analyses further reveal a clear asymmetry: models process code-mixed text through an English-dominant semantic subspace, while native-script Hindi provides complementary signals that reduce representational uncertainty. Motivated by these findings, we introduce a trilingual post-training alignment objective that brings code-mixed representations closer to both constituent languages simultaneously, yielding more balanced cross-lingual alignment and downstream gains on sentiment analysis and hate speech detection -- showing that grounding code-mixed representations in their constituent languages meaningfully helps cross-lingual understanding. Code is available at: https://github.com/debajyotimaz/tri align EMNLP 2026.
Sources
- Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
- CALCS 2021 Shared Task: Machine Translation for Code-Switched Data
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering