Mixture-of-experts routing
DeepSeekMoE
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Heavily superseded — a standard baseline that newer methods routinely beat
0 papers critique it · 6 beat it on benchmarks
Beaten on benchmarks
Head-to-head results where a newer method reports beating DeepSeekMoE. Values are copied from the source paper's tables — verify against the cited paper.
DS-MoE-6B beats DeepSeekMoE
4603.9 vs 3144.1
BlockFFN beats DeepSeekMoE
71.38 vs 49.27
SPHERE beats DeepSeekMoE
0.45 vs 0.33
average · [MetaWorld DS-MoE CRL]
SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement LearningMistral-MoCE beats DeepSeekMoE
55.03 vs 40.77
Avg · [Cross-model comparison]
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction TuningIDA-MoE beats DeepSeekMoE
43.1 vs 38.4
VizWiz · [StableLM-1.6B + CLIP-336]
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsOmniMoE beats DeepSeekMoE
50.9 vs 50.2
Avg · [6.4B-A1.7B models]
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 1, 2026
- May 24, 2026
- May 11, 2026
- SPHERESPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement LearningMay 6, 2026
- May 6, 2026
- Apr 23, 2026
- Feb 10, 2026
- Feb 9, 2026
- Feb 5, 2026
- GRIP (Geometric Routing Invariance Preservation)GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router ConstraintsJan 23, 2026
- Jan 7, 2026