LLM quantization
QuIP
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Superseded — cited as a baseline and beaten by newer methods
3 papers critique it · 4 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites QuIP as a baseline.
It is a powerful method and achieves 2-bit level quantization, but its computation cost is a little expensive.
“Table tab:ppl2048 also shows the importance of incoherence processing. without fine-tuning or lattice codebooks significantly outperforms OmniQuant and AWQ, which both rely on heuristics to reduce model outliers during quantization.”
“However, these predetermined rotations cannot adapt to specific models.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating QuIP. Values are copied from the source paper's tables — verify against the cited paper.
FrameQuant beats QuIP
8.92 vs 1.42
Top-1 accuracy · [2-bit quantization, ViT-T]
FrameQuant: Flexible Low-Bit Quantization for TransformersSEPTQ beats QuIP
16.20 vs 2998.00
ButterflyQuant beats QuIP
15.40 vs 37.32
WikiText-2 Perplexity · [LLaMA2-7B]
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly TransformsPTQ1.61 beats QuIP
39.20 vs 26.26
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 29, 2026
- LFQLFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMsMay 28, 2026
- ADMM-QADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language ModelsMay 11, 2026
- May 6, 2026
- Apr 11, 2026
- Jan 21, 2026
- Grouped Lattice Vector Quantization (GLVQ)Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionOct 23, 2025
- Sep 28, 2025
- Bi-VLMBi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language ModelsSep 23, 2025
- Sep 18, 2025