Mixture-of-experts routing
HydraLoRA
HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning
Heavily superseded — a standard baseline that newer methods routinely beat
1 papers critique it · 8 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites HydraLoRA as a baseline.
Compared to existing state-of-the-art MoE baselines (Switch Transformer, MoLE, HydraLoRA), HiLoMoE consistently shows superior efficiency and effectiveness. On average, it improves AUC by 0.08% and reduces LogLoss by 0.10% compared to the best competing MoE variant (HydraLoRA). At the same time, HiLoMoE reduces parameter count by an average of 4.04K, which is equivalent to a 21.0% reduction relative to the most parameter-efficient MoE competitor (HydraLoRA).
Beaten on benchmarks
Head-to-head results where a newer method reports beating HydraLoRA. Values are copied from the source paper's tables — verify against the cited paper.
SMoRA beats HydraLoRA
36.66 vs 32.41
Average · [Multi-Domain]
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task LearningLoPE beats HydraLoRA
40.78 vs 37.96
MMLU · [30% noise ratio]
Noise-Robustness Through Noise: A Framework combining Asymmetric LoRA with Poisoning MoEFourierMoE beats HydraLoRA
73.24 vs 69.63
AVG. · [LLaMA-3 8B]
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language ModelsGOAT beats HydraLoRA
60.20 vs 57.39
GSM8K · [Natural Language Generation (NLG)]
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization AlignmentMoORE (L=8) beats HydraLoRA
85.11 vs 83.84
Overall · [CSR-MTL multi-task adaptation]
MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task AdaptationHiLoMoE beats HydraLoRA
0.4341 vs 0.4374
LogLoss · [DIEN + KuaiVideo]
Hierarchical LoRA MoE for Efficient CTR Model ScalingLoRA-SMoE-S(60%) beats HydraLoRA
83.2 vs 82.9
average · [multi-task benchmark]
A Sensitivity-Driven Expert Allocation Method in LoRA-MoE for Efficient Fine-TuningLiMEDoRA beats HydraLoRA
78.12 vs 78.11
Vision Benchmark · [Vision Benchmark]
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- PARAMΔ Integration into Upcycled MoEA Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$Δ$ Integration into Upcycled MoEMay 18, 2026
- MEMIT-like framework for MoEScalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured UpdatesMay 15, 2026
- May 11, 2026
- May 8, 2026
- Apr 28, 2026
- CoGR-MoECoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question AnsweringApr 18, 2026
- Apr 2, 2026
- On Token's DilemmaOn Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language ModelsMar 29, 2026
- Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPsScaling Machine Learning Interatomic Potentials with Mixtures of ExpertsMar 9, 2026
- Mar 5, 2026
- Feb 13, 2026
- PASs-MoEPASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual LearningJan 19, 2026