Mixture-of-experts routing
Pre-gated MoE
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
Superseded baseline#22 of 1,370 most-superseded · first seen Aug 23, 2023
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
4 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Pre-gated MoE as a baseline.
However, they still maintain a layer-wise architecture which constrains batching.
“The majority of existing methods are designed for relatively stable resource environments and generally fail to account for the dynamic heterogeneity of edge network resource states and varying application requirements.”
“as in Pre-gated MoE~hwang2024pre, which introduces pre-gating for parallel loading at potential accuracy cost”
“Despite its ability to accurately predict expert usage and execute prefetching, the structural modifications will complicate direct deployment.”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 4, 2026
- May 19, 2026
- CoX-MoECoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-ExecutionMay 18, 2026
- HodgeCoverHodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-ExpertsMay 13, 2026
- Apr 22, 2026
- Apr 12, 2026
- Alloc-MoEAlloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceApr 9, 2026
- Mar 19, 2026
- Mar 13, 2026
- Mar 12, 2026
- Mar 6, 2026