Mixture-of-experts routing
MoE-Pruner
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
Superseded baseline#20 of 1,370 most-superseded · first seen Oct 15, 2024
Superseded — cited as a baseline and beaten by newer methods
4 papers critique it · 2 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites MoE-Pruner as a baseline.
all existing expert-level approaches suffer catastrophic performance collapse on benchmarks such as GSM8K, HumanEval, MBPP, and MATH.
“requires retraining or fine-tuning and produces static pruning decisions dependent on the calibration datasets and requires knowledge distillation to recover from moderate accuracy degradation.”
“However, these approaches often lead to substantial performance degradation due to the permanent loss of expert knowledge.”
“MoE-$I^2$~moei2024 and MoE-Pruner~pruner2024 partially prune expert weights, but struggle to balance identifying important weights and achieving speedup.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating MoE-Pruner. Values are copied from the source paper's tables — verify against the cited paper.
MoEITS beats MoE-Pruner
62.45 vs 50.47
Average (Av) · [Qwen1.5-2.7B pruning]
MoEITS: A Green AI approach for simplifying MoE-LLMs
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 4, 2026
- May 19, 2026
- CoX-MoECoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-ExecutionMay 18, 2026
- HodgeCoverHodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-ExpertsMay 13, 2026
- Apr 22, 2026
- Apr 12, 2026
- Alloc-MoEAlloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceApr 9, 2026
- Mar 19, 2026
- Mar 13, 2026
- Mar 12, 2026
- Mar 6, 2026