LLM quantization

QuIP

QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Superseded baseline#9 of 80 most-superseded · first seen Jul 25, 2023

Superseded — cited as a baseline and beaten by newer methods

3 papers critique it · 4 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites QuIP as a baseline.

It is a powerful method and achieves 2-bit level quantization, but its computation cost is a little expensive.
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
Table tab:ppl2048 also shows the importance of incoherence processing. without fine-tuning or lattice codebooks significantly outperforms OmniQuant and AWQ, which both rely on heuristics to reduce model outliers during quantization.
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
However, these predetermined rotations cannot adapt to specific models.
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms

Beaten on benchmarks

Head-to-head results where a newer method reports beating QuIP. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.