MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training
Yikai Wang, Chuansai Zhou, Yuhang Zhou, Weiqiang Wu, Cong Wu, Yue Deng, Ben Feng, Mingming Zhu, Beirong Zhou, Zhibin Wang, Sheng Zhong, Chen Tian, Wangze Zhang
State Key Laboratory for Novel Software Technology, Nanjing University · Huawei
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: This paper proposes a multi-view, frequency-aware expert pruning method for constructing low-cost proxy models of Mixture-of-Experts (MoE) large language models (LLMs) to enable low-cost failure
Terminology
Summary
This paper proposes a multi-view, frequency-aware expert pruning method for constructing low-cost proxy models of Mixture-of-Experts (MoE) large language models (LLMs) to enable low-cost failure reproduction and diagnosis during reinforcement learning (RL) post-training.
The paper first characterizes practical RL post-training faults on the Huawei Ascend platform, summarizing representative fault categories including Framework and model-state consistency issues,
Operator implementation and numerical consistency issues,
Expert routing consistency issues,
and Optimization and training stability issues.
From these faults, the authors identify three model-side factors relevant to fault reproduction: routing decisions, expert-utilization patterns, and hidden-state representations.
To preserve these characteristics, the proposed method models expert similarity from three complementary perspectives: Router-parameter distance
(using router weight vectors), Co-activation distance
(quantifying dynamic joint-selection behavior via Top-k routing), and Routed-context prototype distance
(preserving input representation characteristics of hidden states routed to different experts). These distances are normalized and fused into a joint distance, which is used for K-Medoids clustering. From each cluster, the most frequently activated expert is selected as the representative, and the proxy model is reconstructed by copying the MLP parameters and corresponding Router weights of selected experts, preserving the original backbone and Top-k routing mechanism without introducing parameter averaging or additional fine-tuning.
The experimental results demonstrate substantial computational savings. For Qwen3-30B-A3B, the proxy model retaining 48 experts reduces accelerator requirements from 16 to 8 NPUs, with per-step NPU-hours decreasing from 0.573 to 0.255, an approximately 55.5% reduction. For DeepSeek-V3.2, the proxy model requires only 64 NPUs instead of 512, reducing per-step computational cost from 426.7 NPU-hours to 12.8 NPU-hours, an approximately 33.3× reduction.
The proxy models preserve major training dynamics, as shown in experiments on DeepSeek-V4-Flash where the proxy models preserve the major training trends of reward improvement and KL-loss growth.
Under fault conditions, the proxy models reproduce consistent anomaly patterns. For the Rollout LogP Precision Fault,
both the original and proxy models exhibit a shift in the Rollout–Training LogP discrepancy and an increase in KL divergence
with consistent anomaly directions and temporal trends.
For the Actor Update Omission Fault,
both models exhibit the same fault-induced suppression of reward improvement.
Ablation studies show the complete three-view method outperforms pairwise-view combinations, achieving GSM8K accuracies of 48.0% and 86.0% when retaining 48 and 64 experts respectively on Qwen3-30B-A3B. On DeepSeek-V3.2, the proposed method achieves 58.5% accuracy with 16 experts, outperforming frequency-only selection (53.0%), random selection (3.0%), and fixed First-K retention (3.0%).
The authors conclude that the proposed proxy models can serve as efficient surrogates for low-cost fault reproduction, validation, and auxiliary diagnosis before deploying expensive verification on the original models.
Improvements for AI systems
Improvements to AI Systems:
-
Fault-Reproduction Proxy Models for RL Post-Training: Build a lightweight, high-fidelity surrogate of any MoE LLM (e.g., Qwen3, DeepSeek) that preserves routing decisions, expert-utilization patterns, and hidden-state representations. This proxy can be used to rapidly reproduce and diagnose training faults (e.g., LogP precision errors, optimizer update omissions) before running expensive full-model validation, cutting compute by up to 33×.
-
Multi-View Expert Clustering for Model Compression: Implement a K-Medoids clustering algorithm that fuses three distance metrics—router-parameter distance, co-activation distance (Top-k joint selection), and routed-context prototype distance—to select representative experts. This yields proxy models that retain original backbone and Top-k routing without parameter averaging or fine-tuning, enabling lossless-ish compression for debugging and experimentation.
-
Low-Cost Training-Dynamics Monitoring: Use the proxy model as a real-time surrogate to track reward improvement and KL-loss growth during RL post-training, preserving major training trends. This allows continuous monitoring of training health on fewer accelerators (e.g., 8 NPUs instead of 16 for Qwen3-30B-A3B) and early detection of divergence or stagnation.
-
Anomaly-Pattern Consistency Checker: Deploy the proxy to generate expected anomaly signatures (e.g., shifts in Rollout–Training LogP discrepancy, KL divergence spikes, suppression of reward improvement) for known fault classes. Compare these signatures against the original model’s behavior to validate whether a suspected fault is present, enabling fast, low-cost triage.
-
Scalable Fault Injection Testing: Use the proxy to run systematic fault-injection campaigns (e.g., corrupting router weights, altering Top-k selection, injecting numerical noise) to map out failure modes and their early indicators. This improves robustness by pre-identifying vulnerable components in the MoE architecture before full-scale training.
What the Improved AI System Can Do:
-
Reproduce and diagnose RL post-training faults on MoE LLMs using a proxy that is up to 33× cheaper in compute (e.g., 12.8 NPU-hours vs. 426.7 per step for DeepSeek-V3.2).
-
Retain high task accuracy (e.g., 86.0% GSM8K on Qwen3-30B-A3B with 64 experts) while using 50% fewer accelerators.
-
Provide consistent anomaly detection across fault types, with matching temporal trends and directional shifts between proxy and original models.
-
Enable rapid iteration on training stability fixes (e.g., optimizer tuning, routing regularization) without full-model overhead, accelerating development cycles.
Sources
- Concrete Problems in AI Safety
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
- Task-Specific Expert Pruning for Sparse Mixture-of-Experts
- Training Verifiers to Solve Math Word Problems
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
- Distilling the Knowledge in a Neural Network
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
- Mixtral of Experts
- DeepSeek-V3 Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
- Training language models to follow instructions with human feedback
- Defeating the Training-Inference Mismatch via FP16
- You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HybridFlow: A Flexible and Efficient RLHF Framework
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks