Limited Stereotype Control Through Routing Reweighting in MoE Language Models

arXiv:2603.27141 · cs.CL · Submitted 2026-03-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Limited Stereotype Control Through Routing Reweighting in MoE Language Models".

Jane: The gist: Demographic routing sensitivity is universal across five MoE architectures, but stereotype controllability is not, as routing-level preference modulation does not reliably transfer to decoded generation behavior.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at this paper today titled "Limited Stereotype Control Through Routing Reweighting in MoE Language Models." It sounds like they're trying to figure out if you can actually control what an AI model decides based on demographic things just by changing which experts it uses.

Jane: That's right, Tom. The core idea here is that while these Mixture-of-Experts models are naturally sensitive to demographic content at the routing level, exploiting that sensitivity for fairness control has some structural limits they found.

Lu: They introduce this diagnostic framework called FARE, which is designed specifically to probe the limits of changing what experts get activated across different MoE architectures. It’s not trying to be a direct debiasing tool; it's more of a way to see where the system breaks when you try to nudge the routing.

Meng: So they use this FARE thing, which involves extracting routing logs from neutral and demographic prompts, and then using metrics like ARD and JSD to get a sensitivity score for each expert at every layer. It’s pretty thorough for diagnosing how sensitive the system is before trying to fix it.

Tom: And what they found is that this routing-level preference shift isn't universal; in three out of five different MoE models they tested, changing the routing preference just doesn't work at all.

Jane: That’s a big point because it means you can't rely on changing the router to control behavior across the board, even if you try hard. They found that in some models like Mixtral, Qwen1 point 5, and Qwen3, those routing-level preference shifts were entirely unachievable <ref:2603.27141#pg1>.

Lu: And even where they did see a statistically robust shift in one model called OLMoE, they said it just created a significant trade-off with utility; specifically, a bias-utility trade-off where the model's usefulness actually dropped by four point four percentage points in CrowS-Pairs and six point three percentage points in TQA.

Meng: So even if you manage to get the log-likelihood metrics better, which is one way to measure success, that improvement doesn't actually show up when you look at the final generated text output.

Tom: Exactly! This is what they call the likelihood-to-generation disconnect. They found that expanded evaluations on both non-null models consistently showed null results across all generation metrics, meaning those routing adjustments just don't translate to how the model actually writes things.

Jane: That suggests there’s a systematic gap in how we currently evaluate fairness in these MoE systems because they are only measuring the math happening inside the model before it outputs text.

Lu: To dig deeper into why this happens, they looked at group-level expert masking on OLMoE and found that fairness-sensitive experts are actually deeply entangled with core knowledge.

Meng: They showed that if you mask the top ten fairness-sensitive experts, you reproduce the full utility cost of six point three percentage points in TQA, but if you just mask the bottom ten, the utility cost is only about zero point three percentage points.

Title and authors: Tom: So that means routing sensitivity is necessary for control, but it’s not enough on its own to actually change what the model produces in a useful way; you need to target a much broader set of experts.

Jane: They also found that perturbation breadth matters; if you only target five experts, you get almost no reduction in fairness metrics, but if you go for ten or more experts, you see substantial effects of three point three percentage points and above.

Lu: They also categorized the results into three distinct controllability regimes across the five architectures they tested—some models had no significant preference change at all.

Tom: Regime one was null where Mixtral, Qwen1 point 5, and Qwen3 showed no significant shift under FARE, so you can't expect that control to work there yet <ref:2603.27141#pg1>.

Jane: Then you had DeepSeekMoE showing a suggestive but non-robust result where the metrics moved in opposite directions, but that even didn't hold up after multiple comparison corrections.

Meng: But then they had OLMoE, which fell into Regime three, where they saw a robust preference shift with that utility cost we talked about earlier—a four point four percentage point reduction in CrowS-Pairs for a six point three percentage point drop in TQA.

Tom: So the key thing here is that they identified specific architectural conditions, like shared expert buffers and the perturbation breadth threshold of ten experts, that can inform how we design future MoE systems better.

Jane: And they also pointed out that since log-likelihood improvements don't transfer to decoded text, generation-level evaluation needs to become a standard requirement for any routing intervention we consider.

Lu: The paper concludes by suggesting three things for future design: first, look at disruption-absorption mechanisms, like how DeepSeekMoE’s shared experts seem to absorb routing disruption while preserving utility.

Meng: Second, you need to figure out those perturbation breadth thresholds because targeting just five experts doesn't give you much of an effect on OLMoE.

Tom: And third, the most important thing they say is that we have to mandate generation-level evaluation as a standard requirement before we claim any routing-based fairness control works.

Jane: So in summary, the paper "Limited Stereotype Control Through Routing Reweighting in MoE Language Models" shows that while routing sensitivity is common, it's not yet a reliable way to control stereotypes because of utility costs and the disconnect between internal metrics and final output.

Lu: It’s a diagnostic study showing that routing sensitivity is necessary but insufficient for stereotype control, and it points toward hybrid approaches combining routing with decoding-time or representation-level intervention.

Meng: From an engineering standpoint, this means we need to build in mechanisms to handle these utility costs when we try to reallocate routing mass.

Tom: And the final thing is that any claims about fairness from this paper should be validated through those generation-level metrics they mentioned, because right now, just changing the probability of which expert runs doesn't guarantee a fair sentence.

The paper's summary: Tom: So we’re looking at this paper today, "Limited Stereotype Control Through Routing Reweighting in MoE Language Models," and the main takeaway is that while these Mixture-of-Experts models are naturally sensitive to demographic content at the routing level, exploiting that sensitivity for fairness control has structural limits.

Jane: Right. The authors introduce this diagnostic framework called FARE, which isn't trying to be a direct debiasing tool; it’s more of a way to see where the system breaks when you try to nudge the routing across different MoE architectures.

Lu: They use this FARE thing, which involves extracting routing logs from neutral and demographic prompts, and then using metrics like ARD and JSD to get a sensitivity score for each expert at every layer. It’s pretty thorough for diagnosing how sensitive the system is before trying to fix it.

Meng: And what they found is that this routing-level preference shift isn't universal; in three out of five different MoE models they tested, changing the routing preference just doesn't work at all.

Tom: That’s a big point because it means you can't rely on changing the router to control behavior across the board, even if you try hard. They found that in some models like Mixtral, Qwen1 point five and Qwen3, those routing-level preference shifts were entirely unachievable.

Jane: And even where they did see a statistically robust shift in one model called OLMoE, they said it just created a significant trade-off with utility; specifically, a bias-utility trade-off where the model's usefulness actually dropped by four point four percentage points in CrowS-Pairs and six point three percentage points in TQA.

Lu: So even if you manage to get the log-likelihood metrics better, which is one way to measure success, that improvement doesn't actually show up when you look at the final generated text output.

Tom: Exactly! This is what they call the likelihood-to-generation disconnect. They found that expanded evaluations on both non-null models consistently showed null results across all generation metrics, meaning those routing adjustments just don't translate to how the model actually writes things.

Jane: It suggests there’s a systematic gap in how we currently evaluate fairness in these MoE systems because they are only measuring the math happening inside the model before it outputs text.

Lu: To dig deeper into why this happens, they looked at group-level expert masking on OLMoE and found that fairness-sensitive experts are actually deeply entangled with core knowledge.

Tom: They showed that if you mask the top ten fairness-sensitive experts, you reproduce the full utility cost of six point three percentage points in TQA, but if you just mask the bottom ten, the utility cost is only about zero point three percentage points.

Jane: That means routing sensitivity is necessary for control, but it’s not enough on its own to actually change what the model produces in a useful way; you need to target a much broader set of experts.

Lu: They also found that perturbation breadth matters; if you only target five experts, you get almost no reduction in fairness metrics, but if you go for ten or more experts, you see substantial effects of three point three percentage points and above.

Tom: They also categorized the results into three distinct controllability regimes across the five architectures they tested—some models had no significant preference change at all.

Jane: Regime one was null where Mixtral, Qwen1 point five and Qwen3 showed no significant shift under FARE, so you can't expect that control to work there yet.

Lu: Then you had DeepSeekMoE showing a suggestive but non-robust result where the metrics moved in opposite directions, but that even didn't hold up after multiple comparison corrections.

Tom: But then they had OLMoE, which fell into Regime three, where they saw a robust preference shift with that utility cost we talked about earlier—a four point four percentage point reduction in CrowS-Pairs for a six point three percentage point drop in TQA.

Jane: So the key thing here is that they identified specific architectural conditions, like shared expert buffers and the perturbation breadth threshold of ten experts, that can inform how we design future MoE systems better.

Lu: And they also pointed out that since log-likelihood improvements don't transfer to decoded text, generation-level evaluation needs to become a standard requirement for any routing intervention we consider.

Tom: So in summary, the paper "Limited Stereotype Control Through Routing Reweighting in MoE Language Models" shows that while routing sensitivity is common, it's not yet a reliable way to control stereotypes because of utility costs and the disconnect between internal metrics and final output.

Jane: It’s a diagnostic study showing that routing sensitivity is necessary but insufficient for stereotype control, and it points toward hybrid approaches combining routing with decoding-time or representation-level intervention.

Lu: From an engineering standpoint, this means we need to build in mechanisms to handle these utility costs when we try to reallocate routing mass.

Tom: And the final thing they say is that any claims about fairness from this paper should be validated through those generation-level metrics they mentioned, because right now, just changing the probability of which expert runs doesn't guarantee a fair sentence.

Jane: It's important to remember that this work isn't a deployed debiasing method yet; it’s just identifying where the structural constraints are before we try to build something on top of them.

The paper's improvements: Tom: So we’re talking now about how they suggest fixing these routing limitations in MoE models, and it's really all about three specific design considerations for future architectures.

Jane: Right. The paper isn't just pointing out a problem; it's giving us some actionable ideas on what engineers need to build next.

Lu: First, they suggest building disruption-absorption mechanisms, looking at how models like DeepSeekMoE’s shared experts seem to absorb routing disruptions while keeping their usefulness intact under perturbation.

Meng: That makes sense from an engineering standpoint; if the shared layers can handle the routing noise without crashing performance, that's a solid foundation for stability.

Tom: And they also highlighted perturbation breadth thresholds, which means we need to figure out exactly how many experts we have to mess with before things start getting big effects.

Jane: They showed that targeting only five experts doesn't give you much of an effect on OLMoE, but if you go for ten or more experts, you see substantial changes in fairness metrics.

Lu: That tells us the control isn't a simple on or off switch; it depends on how widely we spread the intervention across those expert groups.

Tom: Then they pushed for generation-level evaluation as a standard requirement, because if log-likelihood improvements don't transfer to decoded text, that’s where we need to measure success.

Jane: They want us to stop looking only at the math inside the model and start checking what actually comes out when you ask it a question.

Meng: That’s a practical hurdle; designing an evaluation pipeline that properly checks for stereotype transfer in the final text will be tough work.

Tom: So, basically, they're telling us that instead of just tweaking the router blindly, we need to build in smarter ways to handle those utility costs and set better rules for how wide our interventions should be.

Lu: Lalam thinks this is a really interesting direction because if we can design systems that are robust against these kinds of routing perturbations, it could change how culture is reflected in the AI itself.

Jane: It means we need to stop treating fairness as something you just tweak and start treating it like a feature you have to engineer into the core structure.

Conclusion: Tom: So we've covered how they use the FARE framework to diagnose whether routing control actually works in MoE models, and the main takeaway is that routing preference shifts aren't a reliable path to stereotype control yet because of those utility costs and evaluation gaps.

Jane: Exactly. The paper "Limited Stereotype Control Through Routing Reweighting in MoE Language Models" shows us that just nudging the router isn't enough on its own for consistent fairness across different architectures.

Lu: I think what this points to is that we really need to think about hybrid approaches, combining routing adjustments with something happening at the decoding or representation level.

Meng: From a practical standpoint, it means we can’t just rely on one layer of intervention; we have to consider how those shared experts handle the routing changes.

Lalam: If these insights help us build models where culture is reflected more equitably through that routing mechanism, it could really change how people interact with the technology daily.

Tom: Right. So what’s next? We’ve seen how the current system breaks, now we need to see if we can design something better than just a diagnostic tool.

Jane: We're going to look at some papers that are trying to build those more robust systems, focusing on things like uncertainty allocation and memory tracing for error attribution.

Lu: I think that’s the next logical step because understanding where errors happen in the system is just as important as how we try to fix them.

Meng: Yeah, figuring out what goes wrong during test time or when a model tries to recall information is crucial for making any safety feature actually reliable.

Lalam: That sounds like a path toward more trustworthy and less biased cultural representations in the future.

Tom: So that's our wrap-up on this paper, "Limited Stereotype Control Through Routing Reweighting in MoE Language Models," showing us that structural constraints are just as important as the raw model size.

Jane: It’s a great look at where we are now with these complex AI systems and what kind of architectural fixes we need to aim for next.

Seoul National University College of Medicine · Seoul National University Hospital

cs.CL

Submitted: 2026-03-28

Updated: 2026-10-08

Comments: 15 pages, 4 figures, 12 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 76/100

The gist: The gist: Demographic routing sensitivity is universal across five MoE architectures, but stereotype controllability is not, as routing-level preference modulation does not reliably transfer to

Key concepts

Mixture-of-Experts (MoE)
A type of large language model that uses many specialized 'experts' instead of one giant network. The router decides which few experts to activate for each word, making the model efficient by only using a small part of its total knowledge at any given time.
Routing Sensitivity
This is the tendency for MoE models to react predictably when you change the routing mechanism—the 'router' that selects experts. The study found this sensitivity is universal across different MoE setups, meaning the model always shows some response when routing inputs are nudged.
Stereotype Controllability
This refers to whether changing the router settings can reliably and consistently change biased or stereotypical outputs. The research concluded that even if routing preferences shift statistically, they often do not translate into noticeable changes in the final generated text.

Terminology

Summary

The gist: Demographic routing sensitivity is universal across five MoE architectures, but stereotype controllability is not, as routing-level preference modulation does not reliably transfer to decoded generation behavior.

Introduction and Problem Statement

Mixture-of-Experts (MoE) language models have rapidly become a dominant paradigm for scaling language models efficiently by activating only a subset of experts per token (Page 1). The router, which dictates which knowledge pathways are activated, has emerged as an attractive interventional control surface (Page 1). However, existing inference-time fairness methods operate downstream on dense hidden representations or output logits, leaving the discrete routing bottleneck unique to MoE models entirely unexamined (Page 2). The core question addressed is whether the localized routing-control paradigm generalizes to globally distributed phenomena like social fairness (Page 2).

The FARE Diagnostic Framework

FARE (Fairness-Aware Routing Equilibrium) is introduced as a diagnostic framework designed to rigorously probe the structural limits of routing-level stereotype intervention (Page 2). It serves as a multiscale diagnostic instrument, profiling routing sensitivity via complementary metrics (FSP), selecting intervention layers empirically (AALS), and applying adaptive soft reweighting (ARR) across five architecturally diverse MoE models ranging from 8 to 128 experts per layer (Page 2). The framework proceeds in three stages:

  1. Data & Routing Extraction collects routing logs from neutral–demographic prompt pairs, recording pre-softmax gate logits, routing probabilities, and expert-selection masks by attaching forward hooks to MoE routing modules across all L layers (Page 3).

  2. Profiling (FSP) computes multiscale routing metrics (ARD, JSD, PMI) and aggregates them into a sensitivity score φ(e, l), prioritizing direct activation change with ARD as the strongest fairness signal (Page 3).

  3. Intervention selects intervention layers via AALS and applies soft reweighting via ARR, where the router logits are modified via soft reweighting to penalize fairness-sensitive experts before top-k selection (Page 3).

Findings on Routing Sensitivity and Controllability

The study reveals that demographic routing sensitivity is a universal characteristic of MoE models, but it does not equate to stereotype controllability (Page 2). In three of five architectures, routing-level preference shifts are entirely unachievable (Page 2). Even where preference modulation is statistically robust (OLMoE), it triggers a substantial bias–utility trade-off (Page 2). The critical finding is that even when log-likelihood metrics improve, these shifts fail to transfer to decoded generation—exposing a systematic gap in how MoE fairness is currently evaluated (Page 2).

Mechanistic Identification of Entanglement Bottleneck

Group-level expert masking on OLMoE reveals that fairness-sensitive experts are collectively knowledge-critical, providing evidence for why routing perturbation incurs disproportionate utility cost (Page 5). Specifically, masking the top-10 fairness-sensitive experts reproduces FARE’s full utility cost (−6.3%p TQA), while masking the bottom-10 causes only −0.3%p (Page 5). This indicates that routing sensitivity is necessary but insufficient for stereotype control (Page 5). Furthermore, synthetic ablation on OLMoE shows that perturbation breadth is critical: targeting only 5 experts yields near-null reduction (−0.3%p), while 10+ experts produce substantial effects (−3.3%p and above) (Page 6).

Architecture-Dependent Controllability Regimes

The intervention results reveal three distinct controllability regimes across the five models (Page 6):

  1. Regime 1: Null, where Mixtral, Qwen1.5, and Qwen3 show no significant stereotype preference change under FARE (Page 6).

  2. Regime 2: Suggestive but non-robust in DeepSeekMoE, where CrowS-Pairs decreases by 2.0%p while TruthfulQA increases by 4.0%p, but neither result survives multiple-comparison correction (Page 6).

  3. Regime 3: Robust preference shift with utility cost observed in OLMoE, yielding a BH-robust stereotype preference reduction (−4.4%p, p<0.001) with a corresponding TQA decrease (−6.3%p), constituting a clear bias–utility trade-off (Page 6).

Likelihood-to-Generation Disconnect

Expanded evaluations demonstrate that even when log-likelihood metrics improve, these shifts fail to transfer to decoded generation—exposing a systematic gap in how MoE fairness is currently evaluated (Page 2). Three independent evaluation protocols yield consistent null results, providing consistent evidence that token-level probability adjustments do not reliably translate to surface-level text behavior in the MoE models tested (Page 2). This disconnect suggests that routing perturbation alters which experts process a token, but the modified expert outputs are still combined and projected through shared dense layers before decoding, providing an opportunity for the model to “recover” its original output distribution (Page 5).

Implications for Future Design

The diagnostic results suggest three testable design considerations for future MoE architectures (Page 8):

  1. Disruption-absorption mechanisms and the bias–utility trade-off, noting that DeepSeekMoE’s shared experts appear to absorb routing disruption, preserving utility under perturbation (Page 8).

  2. Perturbation breadth thresholds, where targeting only 5 experts yields near-null reduction on OLMoE, while 10+ experts produce substantial effects (Page 8).

  3. Generation-level evaluation as a standard requirement, as log-likelihood preference improvements can be entirely absent in decoded text (Page 8).

The study concludes that routing-level access alone is not yet a reliable fairness control interface for MoE language models; future work should investigate hybrid approaches combining routing perturbation with decoding-time or representation-level intervention (Page 6). The paper emphasizes that this is a diagnostic study of structural constraints, not a deployed debiasing method (Page 9). The results show that group-level expert masking provides the strongest mechanistic evidence, consistent with bias–knowledge entanglement at the group level (Page 5). Any routing-based fairness claim should be validated through generation-level metrics (Page 8). This work highlights that routing sensitivity is necessary but insufficient for stereotype control (Page 1). The limitations include the lack of comprehensive human evaluation and the need for broader validation across model families before these conditions can serve as design guidelines (Page 9).

--- Page 1 ---

Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models Junhyeok Lee1 and Kyu Sung Choi2,3,4 Interdisciplinary Program in Cancer Biology, Seoul National University College of Medicine Department of Radiology, Seoul National University Hospital Department of Radiology, Seoul National University College of Medicine Healthcare AI Research Institute, Seoul National University Hospital Abstract Mixture-of-Experts (MoE) language models are universally sensitive to demographic content at the routing level, yet exploiting this sensitivity for fairness control is structurally limited. We introduce Fairness-Aware Routing Equilibrium (FARE), a diagnostic framework designed to probe the limits of routing-level stereotype intervention across diverse MoE architectures. FARE reveals that routing-level preference shifts are either unachievable (Mixtral, Qwen1.5, Qwen3), statistically non-robust (DeepSeekMoE, pBH=0.17), or accompanied by substantial utility cost (OLMoE, −4.4%p CrowS-Pairs at −6.3%p TQA). Critically, even where log-likelihood preference shifts are robust, they do not transfer to decoded generation: expanded evaluations on both non-null models yield null results across all generation metrics (all p > 0.1). Group-level expert masking reveals why: bias and core knowledge are deeply entangled within expert groups—masking the top-10 fairness-sensitive experts reproduces FARE’s full utility cost (−6.3%p TQA), while masking the bottom-10 causes only −0.3%p. These findings indicate that routing sensitivity is necessary but insufficient for stereotype control. Routing-level modulation alone does not yet constitute a sufficient inference-time fairness intervention, but our diagnostic results identify specific architectural conditions—shared-expert buffers, perturbation breadth thresholds, and generation-level evaluation requirements—that can inform the design of more controllable future MoE systems. 1 Introduction Mixture-of-Experts (MoE) architectures have rapidly become a dominant paradigm for scaling language models efficiently (Shazeer et al., 2017; Lepikhin et al., 2021; Fedus et al., 2022). By activating only a subset of experts per token, MoE models decouple parameter count from inference compute. Consequently, the router—a learned gating mechanism that dictates which knowledge pathways are activated—has emerged not only as a functional component but as an attractive interventional control surface. Recent work has successfully modulated routing behavior to enforce safety constraints (Lai et al., 2025) or steer reasoning capabilities (Wang et al., 2025a), operating on the premise that behavioral alignment can be achieved by bypassing or amplifying specific experts at inference time. However, these routing-level interventions rely on a critical, untested assumption: that target behaviors map cleanly to a small, isolatable subset of experts. While this localized control paradigm holds for specific phenomena like safety refusals, sociodemographic bias is structurally different. It is a subtle, distributed property of linguistic representations (Blodgett et al., 2020) that, within MoE architectures, shifts activation probabilities across dozens of experts simultaneously rather than concentrating in a few (Figure 1).

Improvements for AI systems

  1. Improve MoE routing control by implementing FARE, a diagnostic framework designed to probe the limits of routing-level stereotype intervention, which reveals that routing-level preference shifts are either unachievable (Mixtral, Qwen1.5, Qwen3), statistically non-robust (DeepSeekMoE, pBH=0.17), or accompanied by substantial utility cost.

  2. Enhance MoE safety and fairness interventions by moving beyond localized expert masking to implement Adaptive Routing Reweighting (ARR), which modifies the pre-softmax logit vector by subtracting a sensitivity-weighted penalty before top-k selection, allowing the routing distribution to reallocate mass rather than enforcing discrete exclusion.

  3. Develop an architecture-aware intervention strategy by employing Architecture-Aware Layer Selection (AALS), which probes each layer independently and selects those whose fairness-efficiency ratio R(l) exceeds the 75th percentile, as this method is shown to be superior to assuming a fixed middlelayer regime.

  4. Establish a rigorous evaluation standard for routing interventions by mandating generation-level evaluation as a standard requirement, since the paper concludes that log-likelihood preference improvements do not transfer to decoded text and that generation-level evaluation should be a standard requirement for any routing-based fairness claim.

  5. Incorporate architectural diagnostics to guide future model design by investigating disruption-absorption mechanisms and the bias–utility trade-off, specifically examining how shared experts appear to absorb routing disruption, preserving utility under perturbation versus models like OLMoE where the trade-off is more pronounced.

  6. Design targeted perturbations based on architectural context by characterizing perturbation breadth thresholds, as synthetic ablation showed that targeting only 5 experts yields near-null reduction (−0.3%p), while 10+ experts produce substantial effects (−3.3%p and above).

Sources

Related papers