Living systematic review
KV-cache compression
Cutting the memory and bandwidth cost of the transformer key-value cache in long-context LLM inference — token eviction, quantization/low-rank, offload/reuse, and head/layer-adaptive budgeting.
264 papers613 critique receipts2,449 benchmark resultsupdated Jun 18, 2026
Most-superseded baselines
Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.
- 1SnapKV
SnapKV: LLM Knows What You are Looking for Before Generation
51 critique · 71 beaten on benchmarks
- 2H2Oin SnapKV
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
65 critique · 56 beaten on benchmarks
- 3StreamingLLMin SnapKV
Efficient Streaming Language Models with Attention Sinks
43 critique · 44 beaten on benchmarks
- 4PyramidKVin SnapKV
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
21 critique · 29 beaten on benchmarks
- 5KIVI
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
20 critique · 27 beaten on benchmarks
- 6Quest
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
13 critique · 16 beaten on benchmarks
- 7TOVAin SnapKV
Transformers are Multi-State RNNs
6 critique · 14 beaten on benchmarks
- 8AdaKVin SnapKV
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
7 critique · 7 beaten on benchmarks
- 9CaMin SnapKV
6 critique · 6 beaten on benchmarks
- 10KVQuantin KIVI
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
6 critique · 6 beaten on benchmarks
- 11TurboQuantin KIVI
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
7 critique · 4 beaten on benchmarks
- 12Scissorhandsin SnapKV
Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time
8 critique · 3 beaten on benchmarks
The competition
Methods that fight on the same benchmarks cluster into distinct sub-problems.
Palu30 methods
Palu · ThinK · Eigen Attention · PagedAttention · Loki · Lexico
MiniCache23 methods
MiniCache · CacheBlend · TurboRAG · EPIC · Mooncake · PromptCache
Fast-dLLM11 methods
Fast-dLLM · dKV-Cache · dLLM-Cache · Block diffusion · Elastic-Cache · fixed-schedule KV caching
KVFlow6 methods
KVFlow · CachedAttention · GPU decompression (CacheGen) · Host CPU decompression · PBKV · ShadowServe
Best-of-N6 methods
Best-of-N · Prompted self-correction · tree search · Best-of-16 · Latent Phase-Shift Rollback · Prompted SC
LURE6 methods
LURE · OPERA · simple top-K KV cache pruning · VCD · WoodPecker · PruneHal
The frontier
Recent methods not yet superseded in the knowledge base.
- Jun 7, 2026
- Jun 3, 2026
- Jun 1, 2026
- Jun 1, 2026
- May 31, 2026
- May 30, 2026
- May 29, 2026
- May 28, 2026
- May 28, 2026
- May 26, 2026
- May 26, 2026
- May 25, 2026