Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
Least supersededMost
What papers say
Verbatim critique sentences, each from a paper that cites GPU decompression (CacheGen) as a baseline.
“
Unfortunately, this introduces a new problem: interference with model computation. We find that concurrently running decompression and LLM serving on the same GPU causes both tasks to slow down, often by ~30%.