LLM quantization

HAQ

HAQ: Hardware-Aware Automated Quantization with Mixed Precision

Superseded baseline#73 of 80 most-superseded · first seen Nov 21, 2018

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites HAQ as a baseline.

prohibitively expensive for multi-billion-parameter LLMs due to the computational intensity of second-order matrix computations or extensive policy evaluations
SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.