MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

arXiv:2608.10823 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Yikai Wang, Chuansai Zhou, Yuhang Zhou, Weiqiang Wu, Cong Wu, Yue Deng, Ben Feng, Mingming Zhu, Beirong Zhou, Zhibin Wang, Sheng Zhong, Chen Tian, Wangze Zhang

State Key Laboratory for Novel Software Technology, Nanjing University · Huawei

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper proposes a multi-view, frequency-aware expert pruning method for constructing low-cost proxy models of Mixture-of-Experts (MoE) large language models (LLMs) to enable low-cost failure

Terminology

Summary

This paper proposes a multi-view, frequency-aware expert pruning method for constructing low-cost proxy models of Mixture-of-Experts (MoE) large language models (LLMs) to enable low-cost failure reproduction and diagnosis during reinforcement learning (RL) post-training.

The paper first characterizes practical RL post-training faults on the Huawei Ascend platform, summarizing representative fault categories including Framework and model-state consistency issues, Operator implementation and numerical consistency issues, Expert routing consistency issues, and Optimization and training stability issues. From these faults, the authors identify three model-side factors relevant to fault reproduction: routing decisions, expert-utilization patterns, and hidden-state representations.

To preserve these characteristics, the proposed method models expert similarity from three complementary perspectives: Router-parameter distance (using router weight vectors), Co-activation distance (quantifying dynamic joint-selection behavior via Top-k routing), and Routed-context prototype distance (preserving input representation characteristics of hidden states routed to different experts). These distances are normalized and fused into a joint distance, which is used for K-Medoids clustering. From each cluster, the most frequently activated expert is selected as the representative, and the proxy model is reconstructed by copying the MLP parameters and corresponding Router weights of selected experts, preserving the original backbone and Top-k routing mechanism without introducing parameter averaging or additional fine-tuning.

The experimental results demonstrate substantial computational savings. For Qwen3-30B-A3B, the proxy model retaining 48 experts reduces accelerator requirements from 16 to 8 NPUs, with per-step NPU-hours decreasing from 0.573 to 0.255, an approximately 55.5% reduction. For DeepSeek-V3.2, the proxy model requires only 64 NPUs instead of 512, reducing per-step computational cost from 426.7 NPU-hours to 12.8 NPU-hours, an approximately 33.3× reduction.

The proxy models preserve major training dynamics, as shown in experiments on DeepSeek-V4-Flash where the proxy models preserve the major training trends of reward improvement and KL-loss growth. Under fault conditions, the proxy models reproduce consistent anomaly patterns. For the Rollout LogP Precision Fault, both the original and proxy models exhibit a shift in the Rollout–Training LogP discrepancy and an increase in KL divergence with consistent anomaly directions and temporal trends. For the Actor Update Omission Fault, both models exhibit the same fault-induced suppression of reward improvement.

Ablation studies show the complete three-view method outperforms pairwise-view combinations, achieving GSM8K accuracies of 48.0% and 86.0% when retaining 48 and 64 experts respectively on Qwen3-30B-A3B. On DeepSeek-V3.2, the proposed method achieves 58.5% accuracy with 16 experts, outperforming frequency-only selection (53.0%), random selection (3.0%), and fixed First-K retention (3.0%).

The authors conclude that the proposed proxy models can serve as efficient surrogates for low-cost fault reproduction, validation, and auxiliary diagnosis before deploying expensive verification on the original models.

Improvements for AI systems

Improvements to AI Systems:

  1. Fault-Reproduction Proxy Models for RL Post-Training: Build a lightweight, high-fidelity surrogate of any MoE LLM (e.g., Qwen3, DeepSeek) that preserves routing decisions, expert-utilization patterns, and hidden-state representations. This proxy can be used to rapidly reproduce and diagnose training faults (e.g., LogP precision errors, optimizer update omissions) before running expensive full-model validation, cutting compute by up to 33×.

  2. Multi-View Expert Clustering for Model Compression: Implement a K-Medoids clustering algorithm that fuses three distance metrics—router-parameter distance, co-activation distance (Top-k joint selection), and routed-context prototype distance—to select representative experts. This yields proxy models that retain original backbone and Top-k routing without parameter averaging or fine-tuning, enabling lossless-ish compression for debugging and experimentation.

  3. Low-Cost Training-Dynamics Monitoring: Use the proxy model as a real-time surrogate to track reward improvement and KL-loss growth during RL post-training, preserving major training trends. This allows continuous monitoring of training health on fewer accelerators (e.g., 8 NPUs instead of 16 for Qwen3-30B-A3B) and early detection of divergence or stagnation.

  4. Anomaly-Pattern Consistency Checker: Deploy the proxy to generate expected anomaly signatures (e.g., shifts in Rollout–Training LogP discrepancy, KL divergence spikes, suppression of reward improvement) for known fault classes. Compare these signatures against the original model’s behavior to validate whether a suspected fault is present, enabling fast, low-cost triage.

  5. Scalable Fault Injection Testing: Use the proxy to run systematic fault-injection campaigns (e.g., corrupting router weights, altering Top-k selection, injecting numerical noise) to map out failure modes and their early indicators. This improves robustness by pre-identifying vulnerable components in the MoE architecture before full-scale training.

What the Improved AI System Can Do:

  • Reproduce and diagnose RL post-training faults on MoE LLMs using a proxy that is up to 33× cheaper in compute (e.g., 12.8 NPU-hours vs. 426.7 per step for DeepSeek-V3.2).

  • Retain high task accuracy (e.g., 86.0% GSM8K on Qwen3-30B-A3B with 64 experts) while using 50% fewer accelerators.

  • Provide consistent anomaly detection across fault types, with matching temporal trends and directional shifts between proxy and original models.

  • Enable rapid iteration on training stability fixes (e.g., optimizer tuning, routing regularization) without full-model overhead, accelerating development cycles.

Sources

Related papers