Mixture-of-experts routing
BTX
Superseded baseline#44 of 1,370 most-superseded
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 5 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites BTX as a baseline.
However, unlike our method, both approaches only upcycle the FFN part of the dense (seed or specialized) models.
“This could constitute a limitation in certain settings where such finetuning is unfeasible, for instance, because it requires to aggregate domain data into a single centralized node to train the final MoE model, which could raise concerns about privacy, or simply because of computational costs.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating BTX. Values are copied from the source paper's tables — verify against the cited paper.
DU (r=0.5) beats BTX
19.7 vs 18.5
Avg · [Dense 152M → MoE 8×152M]
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- MetaMoEMetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts UnificationMay 14, 2026
- Apr 20, 2026
- BERT-MoE FrameworkAspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User ReviewsFeb 13, 2026
- null experts within token-choice MoEImproving MoE Compute Efficiency by Composing Weight and Data SparsityJan 21, 2026
- MixtureKitMixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts ModelsDec 13, 2025
- ERMoEERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable SpecializationNov 14, 2025
- Dirichlet-Prior Shaping Loss (DPSL)Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEsOct 1, 2025
- Symphony-MoESymphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsSep 23, 2025