Living systematic review

Mixture-of-experts routing

Scaling LLM capacity with sparsely-activated experts — routing, load balancing, and fine-grained expert design.

655 papers1,456 critique receipts4,254 benchmark resultsupdated Jun 23, 2026

Most-superseded baselines

Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.

  1. 1
    Switch Transformer

    Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

    5 critique · 7 beaten on benchmarks

  2. 2
    MC-SMoE

    8 critique · 12 beaten on benchmarks

  3. 3
    HydraLoRA

    HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning

    1 critique · 8 beaten on benchmarks

  4. 4
    DeepSeekMoE

    DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

    0 critique · 6 beaten on benchmarks

  5. 5
    GShardin Switch Transformer

    GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

    2 critique · 2 beaten on benchmarks

  6. 6
    Soft MoEin DeepSeekMoE

    From Sparse to Soft Mixtures of Experts

    3 critique · 4 beaten on benchmarks

  7. 7
    Tutelin Switch Transformer

    Tutel: Adaptive Mixture-of-Experts at Scale

    6 critique · 1 beaten on benchmarks

  8. 8
    Fiddlerin MC-SMoE

    Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

    7 critique · 2 beaten on benchmarks

  9. 9
    DeepSpeed-MoEin Switch Transformer

    DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

    4 critique · 1 beaten on benchmarks

  10. 10
    Mixtral

    Mixtral of Experts

    1 critique · 2 beaten on benchmarks

  11. 11
    LoRAMoEin HydraLoRA

    LoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin

    3 critique · 5 beaten on benchmarks

  12. 12
    Upcycling

    Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

    3 critique · 2 beaten on benchmarks

The competition

Methods that fight on the same benchmarks cluster into distinct sub-problems.

MC-SMoE220 methods

MC-SMoE · Fiddler · EdgeMoE · MoE-Pruner · Pre-gated MoE · SEER-MoE

HydraLoRA145 methods

HydraLoRA · LoRAMoE · MoELoRA · MixLoRA · MoLE · LEMoE

DeepSeekMoE149 methods

DeepSeekMoE · Soft MoE · StableMoE · ReMoE · Lory · SEUF

Switch Transformer143 methods

Switch Transformer · GShard · Tutel · DeepSpeed-MoE · X-MoE · FlexMoE

Upcycling80 methods

Upcycling · Branch-Train-Merge · CLIP-MoE · BTX · FlexOlmo · BTM

LLaMA-MoE60 methods

LLaMA-MoE · AdaMoE · DISP-LLM · ShortGPT · CMoE · static pruning

Mixtral48 methods

Mixtral · SteerMoE · SafeX · GSPO · GRPO · DPO

MoE-LLaVA34 methods

MoE-LLaVA · Expert Choice · ST-MoE · Loss-free balancing · auxiliary losses · MoCLE

MoEQuant30 methods

MoEQuant · PMQ · Hessian · ODP · EAQuant · uniform bit-width quantization

U-Mamba28 methods

U-Mamba · Vision Transformers · DeblurGAN-v2 · GLARE · Retinexformer · VQCNIR

The frontier

Recent methods not yet superseded in the knowledge base.