Mixture-of-experts routing
SEER-MoE
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
Superseded baseline#23 of 1,370 most-superseded · first seen Apr 7, 2024
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 3 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites SEER-MoE as a baseline.
SEER-MoE removes experts based on activation frequency and fine-tunes with entropy-based regularization.
“Our Sub-MoE explores the merging paradigm that requires neither searching nor fine-tuning.”
“expert pruning metrics based on gate statistics collected during decoding. Although these methods actively deal with expert pruning for MoE models, they are still limited to the machine translation domain with linguistic models.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating SEER-MoE. Values are copied from the source paper's tables — verify against the cited paper.
STUN beats SEER-MoE
70.7 vs 56.7
Avg · [Mixtral-8x7B 25% sparsity]
STUN: Structured-Then-Unstructured Pruning for Scalable MoE PruningHFedMoE beats SEER-MoE
0.444 vs 0.371
test accuracy · [DeepSeek-MoE-16B on MMLU, 30.5 GB]
HFedMoE: Resource-aware Heterogeneous Federated Learning with Mixture-of-Experts
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 4, 2026
- May 19, 2026
- CoX-MoECoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-ExecutionMay 18, 2026
- HodgeCoverHodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-ExpertsMay 13, 2026
- Apr 22, 2026
- Apr 12, 2026
- Alloc-MoEAlloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceApr 9, 2026
- Mar 19, 2026
- Mar 13, 2026
- Mar 12, 2026
- Mar 6, 2026