LLM quantization
SpinQuant
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 8 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites SpinQuant as a baseline.
QuaRot, SpinQuant, and ButterflyQuant do not engage with directly [the regime of per-head q_norm/RoPE compatibility failures]
“previous gradient-based optimization~spinquant cannot easily explore permutation invariance, as permutation creates symmetric local optima in a non-convex fashion.”
“Unlike other learnable methods (e.g., SpinQuant liu2024spinquant) that optimize over the full Stiefel manifold with high computational cost, our sparse parameterization guarantees orthogonality by construction, enabling stable and efficient optimization.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating SpinQuant. Values are copied from the source paper's tables — verify against the cited paper.
SpinQuant + GBS beats SpinQuant
52.84 vs 27.16
ButterflyQuant beats SpinQuant
10.24 vs 17.10
WikiText-2 Perplexity · [LLaMA2-13B]
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly TransformsHeRo-Q beats SpinQuant
10.29 vs 12.14
Wiki2 (Perplexity) · [Llama-3-8B (W4A4)]
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian ConditioningConQuR beats SpinQuant
8.29 vs 9.31
PPL · [Llama-3 70B, 4-4-16]
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMsParetoQ beats SpinQuant
57.2 vs 51.9
Avg. · [3-bit quantization on LLaMA-1B]
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationOffQ beats SpinQuant
6.98 vs 7.4
PPL · [Llama 3-8B]
OffQ: Taming Structured Outliers in LLM Quantization by OffsettingInfoQuant beats SpinQuant
7.07 vs 7.28
WikiText2 perplexity · [4-4-16]
InfoQuant: Shaping Activation Distributions for Low-Bit LLM QuantizationSpecQuant beats SpinQuant
64.75 vs 64.10
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 5, 2026
- FAIR-CalibFAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language ModelsJun 4, 2026
- May 25, 2026
- May 11, 2026
- Activation Residual Hessian Quantization (ARHQ)Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM QuantizationApr 30, 2026
- Apr 20, 2026
- Apr 14, 2026
- Mar 26, 2026
- Jan 29, 2026
- Reasoning-QATWhat Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic StudyJan 21, 2026
- Dec 3, 2025