Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
cs.CL, cs.AI
Submitted: 2026-10-01
Updated: 2026-10-02
Code: https://github.com/lucasbandarkar/xl_moe_
Terminology
Sources
- Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
- Multilinguality in Hybrid Attention LLMs
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- DeepSeek-V3 Technical Report
- MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
- Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
- Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation
- Bactrian-X: Multilingual Replicable Instruction-Following Models with Low-Rank Adaptation
- Language-Specific Latent Process Hinders Cross-Lingual Performance
- Improving Multilingual Language Models by Aligning Representations through Steering
- MTet: Multi-domain Translation for English and Vietnamese
- gpt-oss-120b & gpt-oss-20b Model Card
- Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
- Representation Learning with Contrastive Predictive Coding
- Multilingual E5 Text Embeddings: A Technical Report
- Emu3: Next-Token Prediction is All You Need
- CLEAR: Contrastive Learning for Sentence Representation
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
- Qwen3 Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering