Mixture-of-experts routing
FlexMoE
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
Superseded baseline#33 of 1,370 most-superseded · first seen Apr 8, 2023
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites FlexMoE as a baseline.
each re-balancing introduces significant overhead due to copying optimizer state, limiting the frequency that rebalancing can be performed and thus the efficacy of its adaptive replication.
“still cannot resolve the additional communication of expert parameters”
Beaten on benchmarks
Head-to-head results where a newer method reports beating FlexMoE. Values are copied from the source paper's tables — verify against the cited paper.
ConfSMoE beats FlexMoE
49.18 vs 35.29
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- ConceptM³oEConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational PathologyMay 23, 2026
- DisagMoEDisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe ParallelismMay 10, 2026
- PiperPiper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid ParallelismMay 6, 2026
- GRACE-MoEGRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE InferenceMay 6, 2026
- Apr 21, 2026
- Feb 12, 2026
- Multi-Head LatentMoE and Head Parallel (HP)Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE ParallelismFeb 4, 2026
- Jan 29, 2026
- Rasterized Steered Mixture of ExpertsRasterized Steered Mixture of Experts for Efficient 2D Image RegressionOct 7, 2025
- Sep 24, 2025