Mixture-of-experts routing

FlexOlmo

FlexOlmo: Open Language Models for Flexible Data Use

Superseded baseline#47 of 1,370 most-superseded · first seen Jul 9, 2025

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 1 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites FlexOlmo as a baseline.

freezing shared (non-FFN) parameters during expert training (as done in shi2025flexolmoopenlanguagemodels) significantly degrades performance in our setting
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
reliance on similarity-based proxy selection often produces redundant and narrowly concentrated proxies, limiting coverage of domain-relevant modes and weakening router supervision
MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.