When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry

arXiv:2610.10616 · cs.LG, cs.CR · Submitted 2026-10-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "When Routing Reveals Membership".

Tom: The gist The router-augmented membership inference attack combines conventional outputside signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to reveal whether an…

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Alright everyone, let's talk about this paper: "When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry." This work introduces a new way to find out if some specific data was used to train or fine-tune a Mixture-of-Experts model.

Jane: It seems the authors are looking at how the routing information, which is generated when the model decides which experts to use during inference, can leak whether an input was part of the fine-tuning set.

Lu: The core idea they propose is a router-augmented membership inference attack that mixes standard output signals with these aggregated routing features and uses a classifier trained on shadow models to make this determination.

Meng: So, what's the big claim here? Does this new signal actually give us better privacy protection than just looking at the model's final answer?

Tom: The main finding they report is that router telemetry consistently improves membership inference over a strong set of output-signal signals. They say it increases the True Positive Rate at a one percent False Positive Rate by two point seven to nine point four percentage points across all nine settings they tested <ref:2610.10616#pg1,9.4 percentage points across all nine settings>.

Jane: That sounds pretty significant because it means this routing data is genuinely useful for spotting that fine-tuning usage, even when we're already using a good set of standard output signals.

Lalam: The paper suggests that this leakage doesn't just come from memorizing specific router choices; instead, the fine-tuning process introduces membership information directly into the hidden representations, and the router exposes a projection of that signal even when its parameters are frozen during fine-tuning.

Tom: That’s an interesting distinction there. So, it’s not about memorizing a specific routing path, but about how the internal structure has changed because of that training.

Jane: Exactly. They show that this leakage isn't just tied to the router itself; it reflects membership information encoded in those hidden representations instead of something specific to the router's memory.

Lu: And they back up this idea by showing that their router projection retains more membership information than a matched-rank random subspace, which is a strong comparison against other generic projections.

Meng: That sounds like it’s offering a way to measure how much of the learned knowledge is related to the training data versus just general model behavior.

Tom: The experiments are pretty thorough, showing this effect across three different MoE architectures and three different types of data domains they tested. They also show this leakage persists even under different fine-tuning methods like LoRA or instruction tuning.

Jane: So, if a user is fine-tuning a model using a specific technique, the way the router behaves during inference still gives away that training fact, regardless of the method used to do it.

Lalam: And they show this holds even when we restrict what the attacker can see; for instance, augmenting output signals with just discrete expert selection representation is enough to expose a comparable membership signal.

Tom: That’s interesting because it means you don't need full access to every single routing probability during inference to get that extra benefit.

Jane: Even when only looking at the top-one telemetry, they found an additional gain in True Positive Rate of zero point zero one seven five over just using the standard output signals alone, and this gain gets bigger as you look at more experts.

Lu: It seems like even if you can't see everything, a small piece of that continuous routing information is enough to pull out that membership signal for the classifier they built.

Meng: From an engineering standpoint, it’s encouraging because it means we don't need to build incredibly complex monitoring systems just to get this type of privacy check working effectively.

Tom: But what does this actually mean for the folks who are building and deploying these large models? What’s the practical implication of finding this leakage persists across so many different regimes?

Jane: It means that whatever fine-tuning process happens, there's a consistent way to probe whether that data was used, and this method seems robust across various deployment scenarios.

Lu: The paper suggests that if you want to build defenses against these attacks, you have to consider not just the model's output but also how the routing mechanism is behaving internally during operation.

Meng: So, for practical implementation, it points toward integrating monitoring of these routing signals into the deployment pipeline as a way to detect potential misuse.

Tom: It really shifts where we think about privacy risks in complex AI systems; it’s not just about the model’s final answer anymore, but its entire operational path.

Jane: The paper "When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry" points to a persistent way to check for data usage in Mixture-of-Experts models.

Lu: It shows that the router telemetry provides membership information that complements a broad range of output-side attacks, and this pattern holds for both the accuracy of the attack and how sensitive it is to changes in evaluation metrics.

Meng: The authors note that they found this leakage doesn't need specific router memorization; it reflects membership information encoded in hidden representations rather than just the router itself having specific data stored.

Tom: That means we’re looking at a structural change in the model's internal workings due to fine-tuning, not just a simple memory lookup.

Jane: So, when we think about protecting training data privacy for these models, this paper suggests monitoring those intermediate routing features could be a useful addition to our toolkit.

Conclusion: Tom: So, we're wrapping up on "When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry." Basically, this paper shows that you can use the model’s routing decisions to figure out if someone used a specific example to fine-tune the AI.

Jane: Right. It uses these internal router signals combined with standard output signals and a classifier trained on shadow models to tell it apart.

Lu: The authors are really showing how this leakage isn't tied to memorizing specific routing paths, but rather how the fine-tuning changes the hidden representations themselves.

Meng: So, what does this actually mean for the people building these massive models in the real world? Is it a big deal for data privacy?

Lalam: It means we can build in checks that look at how those internal routing features are behaving during operation, not just looking at the final answer.

Tom: Exactly. The implication is that we might need to look deeper into the model's operational path when thinking about training data usage.

Jane: They've shown this effect across different MoE setups and even different fine-tuning methods, which makes it a pretty consistent finding for AI developers.

Lu: And they prove that this routing projection captures more membership information than just a random generic projection of the hidden state, which is a strong technical point.

Meng: So if we’re talking about practical deployment, it suggests monitoring those intermediate signals could be a way to catch unauthorized fine-tuning activity.

Lalam: It helps improve the culture around privacy by showing that internal mechanisms can reveal sensitive training data even after the model is deployed.

Tom: It shifts our focus from just checking the output to understanding how the model processes things internally during its operation.

Jane: And it gives us a concrete method for probing whether specific training examples are embedded in those complex models.

Lu: We've got a lot more on this paper, and we need to look at how they handle restricted telemetry next.

Yixin Tan, Jiayang Liu, Lu Sun, Yuke Hu, Zheng Li, Rui Wen

Institute of Science Tokyo

cs.LG, cs.CR

Submitted: 2026-10-07

Updated: 2026-10-07

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

The gist: The gist The router-augmented membership inference attack combines conventional outputside signals with aggregated routing features and applies a membership classifier learned from independently

Key concepts

Router Telemetry
This is operational data derived from internal representations of an MoE model during inference. It summarizes routing behavior by averaging the probability assigned to each expert across all layers and experts, providing a fixed-dimensional signal about how the model chose its path for a given input.
Output-Side Membership Signals
These are signals used in membership inference that combine two sources: signals directly from the target model and reference-based signals comparing the fine-tuned model against its original, public pre-trained checkpoint. This creates a comprehensive view of potential leakage from the output.
Router Projection
This technique uses router telemetry to create a continuous representation of routing behavior. The paper demonstrates that this projection retains more membership information than a standard random subspace projection of the hidden state, suggesting it captures unique structural information introduced by fine-tuning.

Terminology

Summary

The gist The router-augmented membership inference attack combines conventional outputside signals with aggregated routing features and applies a membership classifier learned from independently fine-tuned shadow models to reveal whether an example was used to fine-tune the deployed model.

How it works

  1. The attack augments conventional output-side membership inference with router telemetry exposed by the target MoE model Router telemetry is derived from internal representations that may be exposed as an operational signal during inference The attacker combines output-derived membership signals with a fixed-dimensional representation of its routing behavior

  2. Output-side membership signals are constructed by combining signals from the target model and reference-based signals obtained by comparing the fine-tuned model with the public pretrained checkpoint The resulting output-side representation is denoted as phiout(x) = Φout(Mft, Mpre, x)

  3. Router-telemetry features are summarized by using the average routing probability assigned to each expert across all MoE layers and experts to give the router representation phiR(x) = [rl,e(x)]l,e This continuous representation is used by default throughout the main experiments

  4. The router-augmented membership classifier is learned from independently fine-tuned shadow models constructed using auxiliary data from the same domain For each candidate example x queried against the target model, the membership score is s(x) = g([ϕout(x); ϕR(x)])

Key Findings on Leakage

**- Router telemetry consistently improves membership inference over a strong output-signal ensemble, increasing TPR at 1% FPR by 2.7–9.4 percentage points across all nine settings This effect is observed across three MoE architectures and three data domains The leakage persists across full fine-tuning, frozen-router training, LoRA, and instruction tuning Mechanistic analysis shows that the leakage does not require router-specific memorization Instead, fine-tuning introduces membership information into hidden representations, while the router exposes a projection of this signal even when its parameters are frozen The leakage is not specific to router memorization and reflects membership information encoded in fine-tuned hidden representations rather than router-specific memorization. 5.5 EXPERIMENTS > 8. E MECHANISM ANALYSIS > 8. E.2 FROZEN-ROUTER CONTROL > 8. E.1 ROUTER-SUBSPACE PROJECTION > 8> The router projection therefore retains more membership information than a matched-rank random subspace. 5.5 EXPERIMENTS > 8> All three conditions are near the chance-level AUROC of 0.5 before fine-tuning. After one epoch, the router and random projections reach AUROCs of 0.522 and 0.515, increasing to 0.661 and 0.594 after two epochs Together, these results suggest that the observed leakage does not require router-specific memorization. 8> The router telemetry provides membership information that complements a broad range of output-side MIA baselines This pattern holds for both AUROC and TPR @1% FPR, indicating that the additional membership signal is not specific to the output-signal ensemble or to a particular evaluation metric. 5.3 LEAKAGE PERSISTS ACROSS FINE-TUNING REGIMES > 8> Incorporating router telemetry improves TPR @1% FPR under all four regimes, with gains ranging from 0.0305 under LoRA to 0.0904 with a frozen router. The leakage therefore persists across substantially different adaptation settings In particular, freezing the router does not reduce the additional leakage, with a gain comparable to that under full finetuning. 5.4 LEAKAGE UNDER RESTRICTED TELEMETRY AND SHADOW-MODEL ACCESS > 8> Augmenting Aout with the discrete expert-selection representation improves TPR @1% FPR in all nine model–domain settings, with gains ranging from 0.0283 to 0.1043. This indicates that the additional membership leakage does not depend on access to precise routing probabilities, as discrete expert-selection information alone is sufficient to expose a comparable membership signal. 5.4 LEAKAGE UNDER RESTRICTED TELEMETRY AND SHADOW-MODEL ACCESS > 8> Even top-1 telemetry provides an additional ∆ TPR @1% FPR of 0.0175 over Aout. The gain increases monotonically with top-K, reaching 0.0564 at top-K = 8 and 0.0641 at top-K = 32 Thus, membership leakage remains observable even when only a small fraction of the continuous routing information is available. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> The router projection therefore retains more membership information than a matched-rank random subspace. The router projection therefore retains more membership information than a matched-rank random subspace. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> Instead, fine-tuning introduces membership-related structure in hidden representations, and the router exposes a projection of this information even when its parameters are frozen during fine-tuning. 8> The router projection therefore retains more membership information than a matched-rank random subspace. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> This comparison controls for the amount of information retained solely due to dimensionality and tests whether the router subspace captures more membership information than an equally sized generic projection of the hidden state. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> The router projection therefore retains more membership information than a matched-rank random subspace. 8> This comparison controls for the amount of information retained solely due to dimensionality and tests whether the router subspace captures more membership information than an equally sized generic projection of the hidden state. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> The router projection therefore retains more membership information than a matched-rank random subspace. 8> This comparison controls for the amount of information retained solely due to dimensionality and tests whether the router subspace captures more membership information than an equally sized generic projection of the hidden state. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> The router projection therefore retains more membership information than a matched-rank random subspace. 8> This comparison controls for the amount of information retained solely due to dimensionality and tests whether the router subspace captures more membership information than an equally sized generic projection of the hidden state. 5.5 ROUTER TELEMETRY EXPOSES MEMBERSHIP SIGNALS IN HIDDEN REPRESENTATIONS > 8> The router projection therefore retains more membership information than a matched-rank random subspace. 8> This comparison controls for the amount of information retained solely due to dimensionality and tests whether the router subspace captures more membership information than an equally sized generic projection of the hidden state.

Improvements for AI systems

  1. The improved AI system can be deployed in environments where router telemetry is exposed for monitoring, debugging, or load analysis without significantly increasing membership inference risk. This is achieved by applying post-processing perturbations like Gaussian noise, feature dropout, or quantization to the continuous expert-load representation to reduce the additional membership leakage as the perturbation becomes stronger.

  2. The system can maintain a high degree of operational utility while mitigating privacy risks by using restricted telemetry exposure. The paper demonstrates that discrete expert selection information alone is sufficient to expose a comparable membership signal, allowing for the use of a restricted representation where only expert selections are exposed, which still yields significant TPR @1% FPR improvements.

  3. The system can be hardened against memory-specific leakage by ensuring the privacy risk is not tied to router memorization. The mechanistic analysis shows that "the leakage does not require router-specific memorization: fine-tuning introduces membership information into hidden representations, while the router exposes a projection of this signal even when its parameters are frozen." This suggests that defenses should focus on the fine-tuning process rather than solely on altering the router mechanism.

  4. The system can be robust against limited attacker resources by relying on minimal shadow model access. The results show that using only a single shadow model still improves TPR @1% FPR over Aout by 0.0493, indicating that the router-induced membership signal remains exploitable even with minimal shadow-model resources.

Sources

Related papers