Mixture-of-experts routing
StableMoE
StableMoE: Stable Routing Strategy for Mixture of Experts
Superseded baseline#13 of 1,370 most-superseded · first seen Apr 18, 2022
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
3 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites StableMoE as a baseline.
Even methods that specifically target decoupling, like StableMoE~dai2022stablemoe, roller2021hash, struggle because they attempt to learn the routing structure during the most volatile early phases of training; the "teacher" structure they distill from is itself a byproduct of this early-stage volatility.
“However, such methods constrain or freeze routing, limiting the model's ability to adapt routing decisions as representations evolve.”
“While effective, this approach sacrifices routing adaptability.”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 1, 2026
- May 24, 2026
- May 11, 2026
- SPHERESPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement LearningMay 6, 2026
- May 6, 2026
- Apr 23, 2026
- Feb 10, 2026
- Feb 9, 2026
- Feb 5, 2026
- GRIP (Geometric Routing Invariance Preservation)GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router ConstraintsJan 23, 2026
- Jan 7, 2026