LLM quantization

HAWQ

HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision

Superseded baseline#74 of 80 most-superseded · first seen Apr 29, 2019

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites HAWQ as a baseline.

these approaches are prohibitively expensive for multi-billion-parameter LLMs due to the computational intensity of second-order matrix computations
SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.