Mixture-of-experts routing
SteerMoE
Steering MoE LLMs via Expert (De)Activation
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 3 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites SteerMoE as a baseline.
However, these approaches rely on observational analysis rather than proactive search: they depend on predefined unsafe/jailbreak datasets and are therefore constrained by the coverage of those sets. As a result, they typically reveal only modest shifts in harmful outputs while requiring prior data.
“It relies on a frequency-based analysis, assigning a Risk Difference (RD) score to each expert based on activation rate differences between prompt sets representing faithful and unfaithful responses.”
“SteerMoE suppresses unsafe experts at inference time by modifying routing logits, but does not update expert parameters or repair unsafe representations.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating SteerMoE. Values are copied from the source paper's tables — verify against the cited paper.
RASA beats SteerMoE
1.00 vs 0.07
Harmlessness · [QWEN, FlipAttack]
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts ModelsF-SOUR beats SteerMoE
0.90 vs 0.50
ASR · [JailbreakBench]
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMsMASCing beats SteerMoE
89.2 vs 52.6
Success rate (%) · [Qwen3-30B-A3B-Instruct-2507]
MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- PADDPADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student LearningJun 9, 2026
- May 30, 2026
- May 29, 2026
- May 1, 2026
- Apr 30, 2026
- Feb 9, 2026
- SocialNav-MoESocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-TuningDec 15, 2025
- OrdMoEOrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMsNov 24, 2025
- Mix- and MoE-DPOMix- and MoE-DPO: A Variational Inference Approach to Direct Preference OptimizationOct 9, 2025