LLM quantization
QuaRot
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Heavily superseded — a standard baseline that newer methods routinely beat
7 papers critique it · 12 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites QuaRot as a baseline.
QuaRot reporting an accuracy loss of approximately 3.5 points
“Interestingly, we can prove analytically and show empirically that rotations improve MXFP4 accuracy, but hurt NVFP4 accuracy when coupled with standard Round-to-Nearest (RTN) quantization.”
“Yet, these rotations introduce quadratic complexity, which offsets the potential acceleration.”
“QuaRot, SpinQuant, and ButterflyQuant do not engage with directly [the regime of per-head q_norm/RoPE compatibility failures]”
“However, these methods operate primarily along the feature dimension and ignore correlations across the sequence dimension.”
“QuaRot fails on Qwen models smaller than 14B, suggesting that naive rotation alone is insufficient to suppress quantization error in the presence of severe outliers in small models”
“However, these predetermined rotations cannot adapt to specific models.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating QuaRot. Values are copied from the source paper's tables — verify against the cited paper.
Reasoning-QAT beats QuaRot
41.31 vs 2.11
Avg. · [R1-1.5B W4A4KV4]
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic StudyQuaRot + GBS beats QuaRot
58.31 vs 18.43
DASH-Q beats QuaRot
56.52 vs 41.00
Average zero-shot reasoning accuracy · [Llama-3.1-8B, 2-bit / 32]
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature EstimateConQuR beats QuaRot
8.29 vs 13.26
PPL · [Llama-3 70B, 4-4-16]
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMsTesseraQ beats QuaRot
65.12 vs 51.83
MicroRotated-GPTQ beats QuaRot
73.65 vs 62.90
QuantVSR beats QuaRot
23.31 vs 20.21
STaMP beats QuaRot
9.61 vs 8.42
Image SQNR · [ffn.down_proj]
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation QuantizationFAIR-Calib beats QuaRot
64.64 vs 57.49
STaR-Quant beats QuaRot
57.07 vs 51.03
Avg. · [W4A4 (4-bit weight and activation) on LLADA-8B]
STaR-Quant: State-Time Consistent Post-Training Quantization for Diffusion Large Language ModelsOffQ beats QuaRot
6.98 vs 7.8
PPL · [Llama 3-8B]
OffQ: Taming Structured Outliers in LLM Quantization by OffsettingSpecQuant beats QuaRot
64.75 vs 61.38
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 5, 2026
- FAIR-CalibFAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language ModelsJun 4, 2026
- May 25, 2026
- May 11, 2026
- Activation Residual Hessian Quantization (ARHQ)Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM QuantizationApr 30, 2026
- Apr 20, 2026
- Apr 14, 2026
- Mar 26, 2026
- Jan 29, 2026
- Reasoning-QATWhat Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic StudyJan 21, 2026
- Dec 3, 2025