Mixture-of-experts routing

MoE-Pruner

MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router

Superseded baseline#20 of 1,370 most-superseded · first seen Oct 15, 2024

Superseded — cited as a baseline and beaten by newer methods

4 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites MoE-Pruner as a baseline.

all existing expert-level approaches suffer catastrophic performance collapse on benchmarks such as GSM8K, HumanEval, MBPP, and MATH.
Less is MoE: Trimming Experts in Domain-Specialist Language Models
requires retraining or fine-tuning and produces static pruning decisions dependent on the calibration datasets and requires knowledge distillation to recover from moderate accuracy degradation.
XShare: Collaborative in-Batch Expert Sharing for Faster MoE Inference
However, these approaches often lead to substantial performance degradation due to the permanent loss of expert knowledge.
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
MoE-$I^2$~moei2024 and MoE-Pruner~pruner2024 partially prune expert weights, but struggle to balance identifying important weights and achieving speedup.
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis

Beaten on benchmarks

Head-to-head results where a newer method reports beating MoE-Pruner. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.