KV-cache compression
Block diffusion
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
Superseded baseline#124 of 234 most-superseded · first seen Mar 12, 2025
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Block diffusion as a baseline.
This requires considering the KV-Cache in training, making twice the forward computation in the training and its form is still constrained in the autoregressive formula.
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.