KV-cache compression
RTN
Superseded baseline#39 of 234 most-superseded
Superseded — cited as a baseline and beaten by newer methods
0 papers critique it · 4 beat it on benchmarks
Beaten on benchmarks
Head-to-head results where a newer method reports beating RTN. Values are copied from the source paper's tables — verify against the cited paper.
DecoQuant beats RTN
35.9 vs 2.3
Average · [activations only, W 16-A 2]
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache CompressionMiniCache beats RTN
35.44 vs 4.90
Average · [Llama-2-7B-Chat]
MiniCache: KV Cache Compression in Depth Dimension for Large Language ModelsSKVQ beats RTN
37.50 vs 6.76
Average · [Llama-2-7B-chat, Group-size 128, key-cache 2bit, value-cache 2bit, window-size 128]
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language ModelsWindowQuant beats RTN
48.5 vs 44.4
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- SpectrumKVSpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM ServingJun 7, 2026
- Hurwitz Quaternion Multiplicative Quantization (HQMQ)Hurwitz Quaternion Multiplicative Quantization for KV Cache CompressionMay 26, 2026
- May 18, 2026
- May 18, 2026
- TriAxialKVTriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference TasksMay 16, 2026
- KVServeKVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM ServingMay 13, 2026
- WindowQuantWindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference OptimizationMay 4, 2026
- Apr 21, 2026
- eOptShrinkQeOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and QuantizationApr 6, 2026
- Apr 3, 2026
- Mar 30, 2026
- Mar 29, 2026