Mixture-of-experts routing
LEMoE
LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language Models
Superseded baseline#36 of 1,370 most-superseded · first seen Jun 28, 2024
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites LEMoE as a baseline.
However, LEMoE functions as a parameter-preserving framework that attaches external modules to a frozen backbone (typically dense). It addresses routing consistency within the added adaptor rather than the routing distribution shift of the base model itself.
“Additionally, since the routers in these methods are fine-tuned, even if previous experts are frozen, there can still be modifications to past knowledge.”
“LEMoE is based on MoE, but its greedy routing harms old experts' influence when integrating new ones.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating LEMoE. Values are copied from the source paper's tables — verify against the cited paper.
LiveEdit beats LEMoE
96.26 vs 79.43
Average · [LLaVA (7B), 10 edits, VLKEB]
Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- PARAMΔ Integration into Upcycled MoEA Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$Δ$ Integration into Upcycled MoEMay 18, 2026
- MEMIT-like framework for MoEScalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured UpdatesMay 15, 2026
- May 11, 2026
- May 8, 2026
- Apr 28, 2026
- CoGR-MoECoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question AnsweringApr 18, 2026
- Apr 2, 2026
- On Token's DilemmaOn Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsMar 29, 2026
- Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPsScaling Machine Learning Interatomic Potentials with Mixtures of ExpertsMar 9, 2026
- Mar 5, 2026
- Feb 13, 2026
- PASs-MoEPASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual LearningJan 19, 2026