Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders
summary
The gist
Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally — or whether those
In short
The study investigated how multilingual models represent code-mixed text by analyzing their internal representations. They found that pretraining on code-mixed data improves English–code-mixed alignment but hurts English–Hindi alignment. A new trilingual post-training stage, using a specific objective, was introduced to achieve more balanced cross-lingual alignment across English, Hindi, and code-mixed inputs.
Key concepts
- Code-Mixed Representations
- These are the internal mathematical ways multilingual models store text that mixes two or more languages. The researchers found these representations are often weakly tied to both languages initially, primarily leaning on the English representation space.
- EN–CM Alignment Trade-off
- When models are pretrained on mixed language data, improving alignment between English and code-mixed text comes at the expense of alignment between English and Hindi. This suggests that code-mixed data signals are strongest for the EN–CM pair.
- Trilingual Post-Training Alignment
- This is a proposed method where models are further trained using a new objective to push embeddings of sentences from English, Hindi, and code-mixed text closer together in the shared space. This aims to create more balanced alignment across all three language pairs.
Terminology used across episodes
This episode discusses
- Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders · Paper Radio
- Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
- CALCS 2021 Shared Task: Machine Translation for Code-Switched Data
The paper
Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders · Read on arXiv
Indian Institute of Science Education and Research, Bhopal · Microsoft Corporation
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Neither Here Nor There".
Jane: Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the authors of "Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders." It includes Divyansh Pathak, Prashant Kodali, and Jasabanta Patro from the Indian Institute of Science Education and Research.
Jane: That's a great team with deep expertise in multilingual modeling; it tells us this research has a strong foundation in how these massive models are structured.
Lu: Their work is building on prior studies, looking at how English and Hindi representations align well separately, but then seeing where the code-mixed signals actually connect within that shared space.
Meng: What I find interesting is that they are specifically focusing on the representational dynamics rather than just surface-level performance metrics which is a deeper dive into model behavior.
Lalam: It means they aren't just telling us *what* the model gets right, but *how* it’s thinking about that mixed text at a structural level.
The paper's summary: Tom: To summarize what the paper covers, they first constructed a unified trilingual corpus of parallel English, Hindi in Devanagari script, and Hindi–English code-mixed sentences. This gave them the material to test their hypotheses.
Jane: And the core finding they present is that even when standard models align English and Hindi representations effectively on their own, the code-mixed inputs still show a weak connection to either of those constituent languages internally.
Lu: The key observation here is that code-mixed representations remain loosely anchored to both languages under standard pretraining conditions, which leaves the relationship between the mixed input and its monolingual parts underexplored.
Meng: So, it's not a simple failure of translation; it's a structural issue in how the model maps these different linguistic forms into its embedding space when they are mixed together.
Lalam: This means that simply feeding the model more code-mixed data won't automatically fix this fundamental disconnect; we need to address the representation geometry itself.
The paper's improvements: Tom: The paper suggests a way forward by looking at how continued pretraining on code-mixed data specifically impacts alignment, and they found a trade-off. They noted that while continued pretraining on code-mixed data improves English–code-mixed alignment, it comes at the cost of English–Hindi alignment.
Jane: That trade-off is really telling; it shows that adaptation doesn't always improve everything equally, and it can actually weaken a different language pair in this specific context.
Lu: They also pointed out that the script choice matters quite a bit; when Hindi is Romanized during pretraining, the alignment between Hindi and code-mixed text actually degrades.
Meng: From an engineering standpoint, knowing that script preservation helps alignment suggests we need to be careful about how we prepare our input data for models trained on mixed scripts.
Lalam: This points toward a solution where we can guide the model's attention more precisely, maybe through a specific post-training objective that balances these conflicting alignment signals between English and Hindi.
Conclusion: Tom: So, to wrap up this look at "Neither Here Nor There," the authors conclude that while code-mixed pretraining helps English–code-mixed alignment, it's not a complete solution on its own. They propose an explicit trilingual post-training alignment stage using an objective designed to pull semantically equivalent sentences closer together in the shared embedding space.
Jane: That leads us to their main contribution: introducing the Cross-Lingual Alignment Score, or CLAS, as a metric to measure that balanced cross-lingual alignment across all language pairs.
Lu: This stage is designed to encourage more balanced cross-lingual alignment between English, Hindi, and code-mixed representations while trying to keep monolingual consistency strong.
Meng: For practical deployment, this suggests we can use this explicit alignment objective during the fine-tuning phase to ensure that when we deploy a model for a specific application, like sentiment analysis on code-mixed text, its internal logic is robust across all three languages involved.
Lalam: I think the biggest implication for culture is that these improved representations mean our AI can finally handle the rich, nuanced ways people communicate in real-world social media settings without losing meaning or accuracy in any single language component.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck