Mixture-of-experts routing

Expert Choice

Superseded baseline#51 of 1,370 most-superseded

Superseded — cited as a baseline and beaten by newer methods

3 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Expert Choice as a baseline.

Previous strategies like expert-choice anticipated this, but their routing design limit the assignment flexibility to image spatial regions without considering temporal denoising timestep complexity.
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts
Yet, an unacceptable drawback is that it is not suited to casual language modeling due to the reliance on future tokens for the top-k token selection
AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
The Expert Choice approach suffers from token dropping
Unified Sparse Mixture of Experts

Beaten on benchmarks

Head-to-head results where a newer method reports beating Expert Choice. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.