Mixture-of-experts routing

SEER-MoE

SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts

Superseded baseline#23 of 1,370 most-superseded · first seen Apr 7, 2024

Superseded — cited as a baseline and beaten by newer methods

3 papers critique it · 3 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites SEER-MoE as a baseline.

SEER-MoE removes experts based on activation frequency and fine-tunes with entropy-based regularization.
Less is MoE: Trimming Experts in Domain-Specialist Language Models
Our Sub-MoE explores the merging paradigm that requires neither searching nor fine-tuning.
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
expert pruning metrics based on gate statistics collected during decoding. Although these methods actively deal with expert pruning for MoE models, they are still limited to the machine translation domain with linguistic models.
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models

Beaten on benchmarks

Head-to-head results where a newer method reports beating SEER-MoE. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.