Speculative decoding

SpecVLM

SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning

Superseded baseline#23 of 151 most-superseded · first seen Aug 22, 2025

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites SpecVLM as a baseline.

Their experiments with a small VLM draft model incorporating an image encoder yielded only marginal gains, highlighting the challenge of effectively processing visual information in the draft model due to the high redundancy and computational complexity of image inputs.
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
However, existing SD frameworks are fundamentally constrained by their exact-match rule: a draft token is accepted only if it is identical to the target model's generation.
See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs

Beaten on benchmarks

Head-to-head results where a newer method reports beating SpecVLM. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.