Mixture-of-experts routing

Pre-gated MoE

Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference

Superseded baseline#22 of 1,370 most-superseded · first seen Aug 23, 2023

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

4 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Pre-gated MoE as a baseline.

However, they still maintain a layer-wise architecture which constrains batching.
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
The majority of existing methods are designed for relatively stable resource environments and generally fail to account for the dynamic heterogeneity of edge network resource states and varying application requirements.
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
as in Pre-gated MoE~hwang2024pre, which introduces pre-gating for parallel loading at potential accuracy cost
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
Despite its ability to accurately predict expert usage and execute prefetching, the structural modifications will complicate direct deployment.
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.