Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

arXiv:2608.06400 · cs.AI · Submitted 2026-07-31 · Read on arXiv

Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Soyoung Oh, Isabel Valera, Vera Demberg

Saarland University · University of California San Diego · Bielefeld University · Max Planck Institute for Software Systems · Max Planck Institute for Informatics

cs.AI

Submitted: 2026-07-31

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper addresses the limitation that in sparse Mixture-of-Experts (MoE) reward models, "routing weights only reveal which prompts an expert receives, not how it judges responses, providing only a

Terminology

Summary

The paper addresses the limitation that in sparse Mixture-of-Experts (MoE) reward models, routing weights only reveal which prompts an expert receives, not how it judges responses, providing only a partial account of expert behavior. Because routing weights primarily capture the prompt-level domains in which each expert specializes, rather than the response-level preference dimensions that directly explain the model’s decisions, the authors propose a new method.

To address these limitations, we propose CoCo, a Contribution Contrast-based response-level interpretation method for MoE reward models. CoCo "uses contribution contrast as the interpretation signal and characterizes experts using response pairs with large product of the routing weight and absolute expert score difference, thereby capturing the interaction between where an expert is used and how it distinguishes responses. For an expert k, the signed contribution contrast and magnitude" are defined as:

c k(x, y w, y l) = pi phi,k(x) [r theta k(x, y w) - r theta k(x, y l)]

c k(x, y w, y l) = c k(x, y w, y l)

The authors also introduce CoCo-Based Interpretability Regularization. They adapt existing interpretability regularizers to encourage sparse and diverse contribution contrast patterns, specifically applying local sparsity L ls, global balance L gb, and expert diversity L div to the contribution contrast profiles rather than to routing weights. This lightweight adaptation leaves the MoE architecture unchanged, while aligning the interpretability constraints with the contrastive contribution signal used by CoCo.

The method was evaluated through automatic and human evaluations on two datasets (700K and Reddit) against router-based, score-based, and sparse autoencoder (SAE) alternatives. The evaluation metrics included Interpretation Quality (Fidelity and Redundancy), Decision Faithfulness (Expert–Model Agreement and Expert Removal Flip Rate), and Expert Specialization (Expert Accuracy and Relative Expert Advantage).

The results demonstrate that CoCo produces more coherent, faithful, and specialized interpretations than router-based, score-based, and sparse autoencoder (SAE) alternatives while maintaining competitive reward modeling accuracy. Specifically, "CoCo achieves the strongest overall interpretability among existing interpretable reward models across both datasets, obtaining the highest fidelity, expert–model agreement, and expert accuracy while retaining strong reward modeling accuracy. Furthermore, CoCo provides a more informative account of expert behavior than routing weights or expert score differences alone."

Improvements for AI systems

1. Contribution-Contrast Regularized RLHF Training

  • The Improvement: Integrate CoCo-based interpretability regularization (L ls, L gb, and L div) directly into the training objective of Mixture-of-Experts (MoE) reward models, applying constraints to the contribution contrast signal (c k) rather than the routing weights.

  • What the system can do: It will produce reward models where experts are specialized in specific preference dimensions (e.g., one expert focuses on factual accuracy, another on conciseness, and another on logical flow) rather than just prompt domains (e.g., math or coding). This results in more precise Reinforcement Learning from Human Feedback (RLHF) by providing clearer signals for policy optimization.

2. Granular Response-Level Decision Auditing

  • The Improvement: Implement a CoCo-based diagnostic layer in the inference pipeline of reward models to provide real-time, response-specific rationales.

  • What the system can do: Instead of providing a single scalar reward score, the system can generate a detailed audit trail for every decision. It can explain, for example: This response was penalized because Expert 3 identified a high contrast in 'instruction following' between the winning and losing response, even though the prompt was categorized as 'creative writing'. This allows developers to debug exactly why a model prefers one response over another.

3. Targeted Safety and Alignment Guardrails

  • The Improvement: Use CoCo to identify and isolate safety-specialized experts within an MoE reward model that exhibit high contribution contrast on toxic or biased response pairs.

  • What the system can do: The system can act as a high-fidelity safety monitor. When the contribution contrast of a safety expert exceeds a specific threshold for a generated response, the system can trigger an immediate refusal or flag the output for human review, providing much higher precision than monolithic safety classifiers.

4. Contribution-Aware Model Pruning and Compression

  • The Improvement: Replace traditional routing-frequency-based pruning with contribution-contrast-based pruning for MoE architectures.

  • What the system can do: The system can identify and remove redundant experts—those that are frequently routed to (high routing weight) but contribute very little to the actual distinction between good and bad responses (low contribution contrast). This allows for the creation of much smaller, faster, and more efficient reward models that retain the full decision-making nuance of the original large-scale model.

Abstract

Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts through examples with high routing weights. However, routing weights only reveal which prompts an expert receives, not how it judges responses, providing only a partial account of expert behavior. We therefore propose Co ntribution- Co ntrast (CoCo) response-level interpretation, which faithfully characterizes experts' roles using chosen-rejected response pairs with the largest contribution contrasts, jointly capturing routing and preference behavior. Across automatic and human evaluations, CoCo yields more coherent, faithful, and specialized interpretations than router-based, score-based, and sparse autoencoder-based alternatives while maintaining competitive reward modeling accuracy. To the best of our knowledge, this is the first systematic study of interpretation methods for MoE reward models.

Sources

Related papers