Mixture-of-experts routing
SEUF
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
Superseded baseline#43 of 1,370 most-superseded · first seen Nov 27, 2024
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites SEUF as a baseline.
Consequently, the only prior work on MoE unlearning, SEUF zhang2024seuf, relies on soft regularization penalties and restricted expert updates, which risks incomplete forgetting.
“restricting unlearning to only one expert, as in SEUF, may limit the capacity for unlearning.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating SEUF. Values are copied from the source paper's tables — verify against the cited paper.
TRACE beats SEUF
0.5541 vs 0.4553
MMLU · [DeepSeek-V2-Lite-Chat]
Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 1, 2026
- May 24, 2026
- May 11, 2026
- SPHERESPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement LearningMay 6, 2026
- May 6, 2026
- Apr 23, 2026
- Feb 10, 2026
- Feb 9, 2026
- Feb 5, 2026
- GRIP (Geometric Routing Invariance Preservation)GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router ConstraintsJan 23, 2026
- Jan 7, 2026