KV-cache compression

FastV

Superseded baseline#42 of 234 most-superseded

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites FastV as a baseline.

However, most existing methods depend on query-derived text tokens to compute attention scores and initiate compression, inevitably causing response delays in online scenarios.
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
FastV, though training-free, prunes vision tokens without cross-modality guidance, yielding inconsistent results across models and benchmarks.
Make Your LVLM KV Cache More Lightweight

Beaten on benchmarks

Head-to-head results where a newer method reports beating FastV. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.