Mixture-of-experts routing

MoLE

MoLE : Mixture of Language Experts for Multi-Lingual Automatic Speech Recognition

Superseded baseline#31 of 1,370 most-superseded · first seen Feb 27, 2023

Superseded — cited as a baseline and beaten by newer methods

4 papers critique it · 1 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites MoLE as a baseline.

Mixture-of-Lookup-Experts (MoLE) jie2025mole radically replaces MLPs with parameter-free lookup tables, sacrificing expressive power for efficiency.
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
However, the absence of nonlinear activation on the expert outputs may limit the expressive capacity of the resulting representations.
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
Compared to existing state-of-the-art MoE baselines (Switch Transformer, MoLE, HydraLoRA), HiLoMoE consistently shows superior efficiency and effectiveness.
Hierarchical LoRA MoE for Efficient CTR Model Scaling
because the routing is global, the resulting potential applies the same effective weights to every atom in the simulation domain
Mixture of Experts Framework in Machine Learning Interatomic Potentials for Atomistic Simulations

Beaten on benchmarks

Head-to-head results where a newer method reports beating MoLE. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.