Mixture-of-experts routing

X-MoE

XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection

Superseded baseline#21 of 1,370 most-superseded · first seen Feb 27, 2024

Superseded — cited as a baseline and beaten by newer methods

3 papers critique it · 3 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites X-MoE as a baseline.

pure cosine scoring eliminates magnitude cues, whereas SIPS retains them with bounded influence.
L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts
However, for 500B+ models, X-MoE achieves only 5% MFU.
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
addressed representation collapse by routing in a low-dimensional space, but experts still operated on high-dimensional inputs.
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism

Beaten on benchmarks

Head-to-head results where a newer method reports beating X-MoE. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.