RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models".
Jane: Mixture-of-Experts (MoE) models offer efficient scaling for Large Language Models, but adapting them to non-English downstream tasks remains challenging because existing fine-tuning methods treat MoEs as monolithic learners,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: To get into the specifics, we need to look at who wrote this and what their main argument is in "RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models." The title itself tells us they are focusing on routing alignment for multilingual adaptation, which points directly to leveraging the internal structure of MoE models.
Jane: And the authors include Guanzhi Deng, Kuan Wu, Haibo Wang, Shing Yin Wong, Sichun Luo, and Linqi Song from City University of Hong Kong and Carnegie Mellon University. They bring together expertise from different institutions to tackle this complex scaling issue.
Lu: Their core argument is that existing fine-tuning methods fail because they treat MoE models like monolithic learners, completely ignoring the heterogeneous routing structure that develops during pretraining. They validate across multiple MoE models and downstream tasks that middle layers form a language-universal alignment zone where routing divergence strongly predicts per-language task performance gaps.
Meng: So, they aren't just tweaking weights randomly; they are using the observed routing patterns to diagnose *why* performance drops in nonEnglish tasks, pinpointing that the issue is often related to not engaging the right experts for those specific inputs. That’s a much more structured way to approach model tuning.
Lalam: It means we stop wasting compute trying random adjustments and start targeting the exact mechanisms that are already showing promise in cross-lingual transfer, which is a very smart direction for making our AI more capable globally. This structured approach seems much more scalable than trial and error.
The paper's summary: Tom: Now let's look at what they actually propose as the solution in "RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models." They introduce a three-stage framework called RAMoE, which is designed to exploit that routing structure we talked about earlier.
Jane: The summary explains that the framework categorizes parallel task examples into a four-way taxonomy: cc, ci, ic, and ii based on correctness in English and the target language. This classification helps them isolate where the performance gaps are happening.
Lu: The second stage involves identifying task-relevant experts in the middle layers by contrasting routing weights from English task data against a general English corpus, selecting experts based on a "task-specificity score". This procedure localizes those specific experts within that language-universal zone.
Meng: And the third stage is the Routing-Aligned SFT, where they augment standard cross-entropy loss with a routing alignment loss applied specifically to ci-type examples. This loss makes sure the target language routing on those specific examples follows the English task-expert activation pattern.
Lalam: Essentially, they are using this taxonomy to guide their fine-tuning so that when the model sees an example that is correct in English but wrong in the target language, it learns to route toward the same expert activations it used for its English performance. It’s a very targeted intervention based on empirical routing data.
The paper's improvements: Tom: Regarding the specific improvements they detail in "RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models," the paper highlights several key findings that show how much better this method is compared to standard SFT.
Jane: The most significant finding they point to is that the ci proportion, which represents examples correct in English but incorrect in the target language, serves as a reliable predictor of alignment benefit, showing an r-value of zero point seven zero. This means we know exactly where we are likely to see a performance gain from applying this routing alignment technique.
Lu: They also showed that the task experts identified in the middle layers are largely language-agnostic and transfer well across linguistically distant language pairs without needing to re-run the expert identification stage. That transferability across languages is a huge piece of evidence supporting their methodology.
Meng: The ablation studies confirm that removing the routing alignment loss entirely, reverting back to standard SFT, caused the largest single drop in performance, which confirms that this alignment is a primary driver of improvement. It proves the necessity of this specific adjustment.
Lalam: Plus, they found that restricting the alignment loss only to task experts within those middle layers provided a "cleaner and more informative alignment target". That level of specificity in targeting the intervention is what makes this framework so powerful for practical deployment.
Conclusion: Tom: So to wrap up our discussion on "RA-MoE: Routing-Aligned Fine-Tuning for Multilingual Adaptation of Mixture-of-Experts Models," the authors successfully bridge the gap between understanding routing structure and fine-tuning for multilingual MoE models. They’ve shown a concrete mechanism to boost nonEnglish downstream task performance by aligning target language routing toward English task-expert activation patterns in the middle layers.
Jane: It really confirms that those middle layers are stable, language-universal zones where expertise is already encoded, making them ideal targets for this kind of focused intervention. This approach moves us away from treating the entire model as one entity when adapting it for different languages.
Lu: The implication here is that we can use the existing pretraining structure to build a much more adaptable system, potentially reducing the need for massive language-specific fine-tuning datasets. It opens up possibilities for building systems that naturally handle linguistic variation better.
Meng: From an engineering perspective, it means we can simplify our deployment pipelines by focusing on identifying those task experts in the middle layers once, and then applying this targeted fine-tuning mechanism across many languages later. It's about making the adaptation process more modular.
Lalam: I'm really excited because this work suggests we can build AI that is inherently more flexible, capable of handling diverse linguistic inputs with targeted, efficient adjustments rather than brute-force retraining. It’s a step toward truly versatile AI.
City University of Hong Kong · Carnegie Mellon University
cs.CL, cs.AI
Submitted: 2026-05-27
Updated: 2026-09-30
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: Mixture-of-Experts (MoE) models offer efficient scaling for Large Language Models, but adapting them to non-English downstream tasks remains challenging because existing fine-tuning methods treat
Key concepts
- Mixture-of-Experts (MoE)
- MoE models use multiple specialized neural networks, or 'experts,' to handle different parts of a task. Instead of using one large network, the model learns to route an input to the most appropriate expert for that specific piece of information.
- Routing Divergence
- This measures how much the way a MoE model directs its inputs changes across different layers. The paper observes that middle layers show strong alignment between languages, while early and late layers are language-specific, indicating where universal expertise resides.
- Routing-Aligned SFT
- This is a fine-tuning technique where the standard training loss is modified. It adds an extra loss term that forces the model's routing for non-English inputs to mimic the routing patterns observed in English task data, ensuring it uses the correct task experts.
- ci Proportion
- This refers to the proportion of task-language pairs where both languages are correct or both are incorrect. The paper found this ratio reliably predicts how much benefit is gained from routing alignment, suggesting that performance gaps driven by language differences rather than general knowledge are most amenable to this technique.
Terminology
Summary
Mixture-of-Experts (MoE) models offer efficient scaling for Large Language Models, but adapting them to non-English downstream tasks remains challenging because existing fine-tuning methods treat MoEs as monolithic learners, ignoring their heterogeneous routing structure. This paper introduces RAMoE (Routing-Aligned MoE Fine-Tuning), a three-stage framework that leverages the cross-lingual routing dynamics of pretrained MoEs to improve performance on non-English tasks by aligning target language routing toward English task patterns in middle layers.
The Core Observation and Motivation
The research is grounded in an empirical observation regarding the layer-wise behavior of routing divergence in pretrained MoEs. The authors validate a U-shaped layer-wise pattern in cross-lingual routing divergence,
where early and late layers are language-specific, while the middle layers exhibit strong cross-lingual alignment.
Crucially, they demonstrate that languages with larger middle-layer routing divergence consistently exhibit larger task performance gaps relative to English. This suggests that the middle layers already encode language-universal, task-relevant expertise,
and non-English performance degradation stems partly from failing to engage these experts for non-English inputs.
The Three-Stage Framework (RAMoE)
RA-MoE is a three-stage framework designed to exploit this routing structure:
-
Parallel data construction and categorization: The process involves running inference on parallel task data to partition examples into four categories based on correctness in English and the target language:
cc (both correct), ci (English correct, target incorrect), ic (English incorrect, target correct), and ii (both incorrect).
-
Middle layer and Task Expert Identification: This stage profiles routing distributions at each layer to identify the middle layer range, defined as
the longest contiguous segment of layers whose mean divergence falls below the median of the per-layer divergence distribution.
Within this range, task experts are identified by contrasting routing weights from English task data against a general English corpus, selecting experts based on atask-specificity score
where higher scores indicate a preference for the downstream task. -
Routing-Aligned SFT: Standard cross-entropy loss is augmented with a routing alignment loss applied selectively to ci-type examples. This loss encourages the target language routing on these examples to
follow the English task-expert activation pattern,
specifically by computing the KL divergence between the current target-language distribution and a reference distribution restricted to the identified task experts, ensuringthe gradient signal encourages the target-language routing to move toward the English pattern.
Key Findings and Performance Predictors
The experiments across three MoE models, three tasks (GSM8K, IFEval, MMLU), and six languages demonstrated that RA-MoE consistently outperforms standard SFT and strong baselines like Routing Steering and RISE. Key findings include:
ci proportion of a task-language pair serving as a reliable predictor of alignment benefit.
The authors quantified this relationship, finding that the ci proportion predicts the benefit from routing alignment (r=0.70, p=0.001),
indicating that the gain is most pronounced on pairs where performance gaps are language-driven rather than knowledge-driven.
Furthermore, they showed that task experts identified in middle layers are largely language-agnostic and transfer well across linguistically distant language pairs without re-running the expert identification stage.
Ablation and Robustness
Ablation studies confirmed the necessity of each component. Removing the alignment loss entirely (reverting to SFT) caused the largest single drop,
confirming routing alignment as a primary driver. Restricting alignment to only task experts within middle layers, rather than applying it across all experts in the identified range, provided a cleaner and more informative alignment target.
Sensitivity analysis on hyperparameters showed that RA-MoE is robust across variations in the weight scalar λ and the number of task experts K, with optimal performance achieved when λ=1.0 and K=8. The computational cost analysis showed that Stage 3 is on par with standard SFT, while the preparatory stages (Stage 1+2) are manageable, especially when leveraging cross-lingual transferability of intermediate-layer task experts.
Conclusion
RA-MoE successfully bridges the gap between routing structure and fine-tuning for multilingual MoE models. It provides a mechanism to improve non-English downstream task performance by explicitly aligning target language routing toward English task-expert activation patterns in the middle layers, confirming that middle-layer expertise is a stable, language-universal zone amenable to targeted intervention.
Improvements for AI systems
Based on the scientific paper Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
(RA-MoE), here are the specific improvements you can make to AI systems and what those improved systems can achieve:
) Improvements and Capabilities of RA-MoE
The proposed framework, RAoE (Routing-Aligned MoE Fine-Tuning), fundamentally changes how Mixture-of-Experts (MoE) models adapt to non-English tasks by exploiting their pretraining routing structure. By aligning the model's routing during fine-tuning with the successful English task patterns in specific middle layers, you achieve superior cross-lingual performance gains over standard Supervised Fine-Tuning (SFT).
Here are the specific improvements and capabilities:
Abstract
Mixture-of-Experts (MoE) models enable efficient LLM scaling, yet adapting them to non-English downstream tasks remains challenging. Standard multilingual fine-tuning largely ignores their heterogeneous routing structure. Across multiple MoE models and tasks, we find strong cross-lingual routing alignment in middle layers, with routing divergence associated with target-language performance gaps. Motivated by this observation, we propose RA-MoE (Routing-Aligned MoE Fine-Tuning), a three-stage framework for multilingual MoE adaptation. RA-MoE categorizes parallel examples into four correctness groups (cc/ci/ic/ii) and identifies task-relevant experts in middle layers. It then selectively aligns target-language routing on ci examples toward successful English routing patterns, jointly matching the total routing mass assigned to task experts and its relative allocation among them. Experiments across three MoE models, three downstream tasks, and six target languages show that RA-MoE consistently outperforms standard SFT and strong routing-aware baselines. Further analyses confirm the intended routing changes and reveal that middle-layer task routing is largely shared and transferable across languages, providing mechanistic evidence for the cross-language transferability of task-specific routing.
Sources
- Understanding Multilingualism in Mixture-of-Experts LLMs: Routing Mechanism, Expert Specialization, and Layerwise Steering
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
- Measuring Massive Multitask Language Understanding
- Mixtral of Experts
- NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extension
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- DeepSeek-V3 Technical Report
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering