Mixture-of-experts routing
EdgeMoE
Superseded baseline#16 of 1,370 most-superseded · first seen May 30, 2023
Superseded — cited as a baseline and beaten by newer methods
5 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites EdgeMoE as a baseline.
EdgeMoE's static approach determines optimal bit widths based on specific dataset profiling, leading to inflexibility across diverse environments and potential accuracy impacts.
“However, this method does not match the dynamic activation characteristic of MoE models.”
“The majority of existing methods are designed for relatively stable resource environments and generally fail to account for the dynamic heterogeneity of edge network resource states and varying application requirements.”
“While EdgeMoE demonstrates comparable perplexity to D²MoE, it exhibits reduced accuracy in specific zero-shot tasks. This disparity can be ascribed to the differing significance attributed to the various markers, with EdgeMoE utilizing a predetermined mixture of bit-width to ascertain the importance of the experts, resulting in diminished accuracy.”
“While static mixed-precision (e.g., EdgeMoE) addresses this, it remains oblivious to dynamic input complexity at runtime”
Beaten on benchmarks
Head-to-head results where a newer method reports beating EdgeMoE. Values are copied from the source paper's tables — verify against the cited paper.
D²-MoE beats EdgeMoE
4.09 vs 4.38
Perplexity · [Mixtral 8×7B]
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 4, 2026
- May 19, 2026
- CoX-MoECoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-ExecutionMay 18, 2026
- HodgeCoverHodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-ExpertsMay 13, 2026
- Apr 22, 2026
- Apr 12, 2026
- Alloc-MoEAlloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferenceApr 9, 2026
- Mar 19, 2026
- Mar 13, 2026
- Mar 12, 2026
- Mar 6, 2026