LLM quantization
FlatQuant
FlatQuant: Flatness Matters for LLM Quantization
Superseded baseline#12 of 80 most-superseded · first seen Oct 12, 2024
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 4 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites FlatQuant as a baseline.
FlatQuant flatquant demonstrates a loss of only 1.4 points
“Even though these approaches successfully quantize LLMs to 4 bits with slight performance degradation, they apply the same type of transformation across all layers, ignoring the distribution characteristics of each layer within LLMs.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating FlatQuant. Values are copied from the source paper's tables — verify against the cited paper.
Reasoning-QAT beats FlatQuant
21.44 vs 17.34
Avg. · [Qwen3-0.6B W4A4KV4]
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic StudyFAIR-Calib beats FlatQuant
66.66 vs 63.98
SliderQuant beats FlatQuant
5.71 vs 5.79
WikiText2 · [W4A4]
SliderQuant: Accurate Post-Training Quantization for LLMs
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 5, 2026
- FAIR-CalibFAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language ModelsJun 4, 2026
- May 25, 2026
- May 11, 2026
- Activation Residual Hessian Quantization (ARHQ)Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM QuantizationApr 30, 2026
- Apr 20, 2026
- Apr 14, 2026
- Mar 26, 2026
- Jan 29, 2026
- Reasoning-QATWhat Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic StudyJan 21, 2026
- Dec 3, 2025