Mixture-of-experts routing
uniform bit-width quantization
Superseded baseline#241 of 1,370 most-superseded
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
2 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites uniform bit-width quantization as a baseline.
such approaches overlook the sparsity inherent to the MoE architecture, leading to suboptimal performance
“vanilla uniform bit-width quantization and expert pruning based solely on routing scores struggle to maintain performance at extremely high compression ratios”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 22, 2026
- May 21, 2026
- KBVQ-MoEKBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language ModelsJan 30, 2026
- Oct 13, 2025