Mixture-of-experts routing

MoCLE

Superseded baseline#223 of 1,370 most-superseded

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

2 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites MoCLE as a baseline.

However, MoCLE clusters different instruction sets and distributes them to different experts, which compromises the flexibility and autonomy of the experts. Differently, MoE-LLaVA relies on knowledge-rich routers.
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
even though samples have similar instruction embeddings, they may generate distinct parameter optimization directions due to distinct targets, so embedding-based routing is still risky for optimization interference within an expert.
Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.