LLM quantization
OmniQuant
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Heavily superseded — a standard baseline that newer methods routinely beat
2 papers critique it · 14 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites OmniQuant as a baseline.
while AWQ falls apart at even 2.15 bits omniquant and OmniQuant produces unusable models at 2 bits, produces high quality models that are close to OmniQuant 3 bit models.
“However, these methods cannot effectively quantize the LLMs to 4-bit weights and activations.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating OmniQuant. Values are copied from the source paper's tables — verify against the cited paper.
LieQ beats OmniQuant
71.84 vs 23.75
Layer-Wise High-Impact Parameter Ratio Optimization beats OmniQuant
14.64 vs 90.64
Perplexity · [LLaMA-2-7B, W2A16]
Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language ModelsTesseraQ beats OmniQuant
8.05 vs 37.37
Perplexity · [W2A16]
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block ReconstructionSEPTQ beats OmniQuant
9.93 vs 30.84
ButterflyQuant beats OmniQuant
15.40 vs 37.32
WikiText-2 Perplexity · [LLaMA2-7B]
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly TransformsQuaRot beats OmniQuant
6.10 vs 14.26
PPL · [4-bit, Llama 7B]
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsQuIP# beats OmniQuant
6.12 vs 12.28
PTQ1.61 beats OmniQuant
37.20 vs 25.35
SliderQuant beats OmniQuant
8.34 vs 14.26
SignRoundV2 beats OmniQuant
57.88 vs 46.98
Avg · [Llama2-7B, 2 bits, Group Size 128]
SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMsParetoQ beats OmniQuant
61.3 vs 57.1
Avg. · [3-bit quantization on MobileLLM-1B]
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM QuantizationOmniQuant+LFQ beats OmniQuant
79.76 vs 78.17
GSM8K (greedy) · [Llama 3.1 8B, W4, OmniQuant]
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 29, 2026
- LFQLFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMsMay 28, 2026
- ADMM-QADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language ModelsMay 11, 2026
- May 6, 2026
- Apr 11, 2026
- Jan 21, 2026
- Grouped Lattice Vector Quantization (GLVQ)Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionOct 23, 2025
- Sep 28, 2025
- Bi-VLMBi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language ModelsSep 23, 2025
- Sep 18, 2025