Mixture-of-experts routing

StableMoE

StableMoE: Stable Routing Strategy for Mixture of Experts

Superseded baseline#13 of 1,370 most-superseded · first seen Apr 18, 2022

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

3 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites StableMoE as a baseline.

Even methods that specifically target decoupling, like StableMoE~dai2022stablemoe, roller2021hash, struggle because they attempt to learn the routing structure during the most volatile early phases of training; the "teacher" structure they distill from is itself a byproduct of this early-stage volatility.
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
However, such methods constrain or freeze routing, limiting the model's ability to adapt routing decisions as representations evolve.
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
While effective, this approach sacrifices routing adaptability.
Spectral Manifold Regularization for Stable and Modular Routing in Deep MoE Architectures

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.