RAZOR: Pruning Replaceable Experts in LLMs
cs.LG, cs.CL
Submitted: 2026-09-24
Updated: 2026-09-30
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- It Takes a MAESTRO To Prune Bad Experts
- Less is MoE: Trimming Experts in Domain-Specialist Language Models
- Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
- Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
- Half the Experts, All the Code: One-Shot Domain Pruning of Mixture-of-Experts LLMs for Coding
- REAM: Merging Improves Pruning of Experts in LLMs
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
- AIMER: Calibration-Free Task-Agnostic MoE Expert Pruning
- How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle
- Pointer Sentinel Mixture Models
- SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Olmo 3
- MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
- ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
- SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs
- Does a Global Perspective Help Prune Sparse MoEs Elegantly?
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks