KV-cache compression
TOVA
Transformers are Multi-State RNNs
Superseded — cited as a baseline and beaten by newer methods
6 papers critique it · 14 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites TOVA as a baseline.
these methods require access to the full attention matrix, making them incompatible with Flash Attention~flashattention and thus impractical for modern deployment scenarios
“They fix the budget of KV Cache in a finite level, but don't distinguish the differences between layers and between heads.”
“These methods, however, often overlook the structure of key information distribution by naively evicting tokens across the entire sequence.”
“While effective, most methods either discard unused tokens too early or require full cache for scoring.”
“However, these methods rely primarily on attention weights and often overlook the contribution of value states in shaping the final model outputs.”
“TOVA~oren2024tova retains attention sinks and a sliding window of recent tokens; a credential at relative depth 0.5 sits 2,000 tokens outside the window.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating TOVA. Values are copied from the source paper's tables — verify against the cited paper.
KVTC beats TOVA
99.3 vs 0.3
LITM · [MN-Minitron 8B]
KV Cache Transform Coding for Compact Storage in LLM InferenceCompactor beats TOVA
38.8 vs 30.9
Total · [Llama 3.1 at 10% KV retention]
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage ScoresCAKE beats TOVA
29.29 vs 25.31
Avg. · [Llama2-7B-Chat, $B_{\text{total}}=128L$]
CAKE: Cascading and Adaptive KV Cache Eviction with Layer PreferencesDMS beats TOVA
53.3 vs 46.7
accuracy · [CR4, AIME 24, 7B]
Inference-Time Hyper-Scaling with KV Cache CompressionAhaKV beats TOVA
41.84 vs 37.99
KeyDiff beats TOVA
49.19 vs 45.31
Avg. · [Llama3.1-8B at 6K cache budget]
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained EnvironmentsWeightedKV beats TOVA
4.30 vs 4.69
Perplexity · [Cache Size = 256]
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language ModelsExpected Attention beats TOVA
50.25 vs 48.14
score · [Qwen, Longbench, 25% compression]
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries DistributionTA beats TOVA
30.4 vs 31.2
Perplexity · [K=16]
Transactional Attention: Semantic Sponsorship for KV-Cache RetentionOBC+TOVA beats TOVA
45.95 vs 45.01
Avg. · [10% KV (q-Aware)]
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM InferenceTOVA+CAOTE beats TOVA
38.08 vs 37.52
Avg accuracy · [Llama 3.1-8B, 2k budget]
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token EvictionTA-Rgx beats TOVA
100 vs 0
Accuracy · [Llama-3.2-1B, K=16]
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- STaR-KVSTaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language ModelsJun 1, 2026
- May 29, 2026
- May 28, 2026
- May 26, 2026
- May 25, 2026
- CONF-KVCONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLMMay 24, 2026
- May 21, 2026
- May 12, 2026
- Global Retention-Based KV EvictionMake Each Token Count: Towards Improving Long-Context Performance with KV Cache EvictionMay 10, 2026
- ReST-KVReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal SmoothingMay 9, 2026
- May 8, 2026