This paper systematically studies expert-selection redundancy and dynamic expert pruning in twelve fine-grained mixture-of-experts language-model checkpoints from nine architecture families. Uniform top-k truncation is evaluated across knowledge QA, mathematics, code generation, and general reasoning, then compared at matched average routed-expert budgets with four adaptive allocation rules. The authors report that keeping roughly two thirds of selected experts largely preserves aggregate performance and increases inference throughput, while adaptive routing contributes little at conservative budgets but helps under aggressive pruning, especially on generative tasks. Matched comparisons further associate greater pruning resilience with larger and thinking models, and greater vulnerability with a multimodal model whose routing weights are less concentrated.
The study suggests that much fine-grained MoE inference cost can be removed with a simple uniform reduction before sophisticated token-adaptive routing is needed. Its evidence also warns that short-answer QA can hide failures in long-form generation and that safe pruning levels depend on model type. The conclusions are empirical for the reported checkpoints, tasks, implementations, and budgets; they do not establish that the same operating points or rule rankings generalize to every MoE model or serving stack.
The paper’s claims are available in Research claims.