Speculative decoding
Speculative Sampling
Speculative Sampling for Parametric Temporal Point Processes
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 10 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Speculative Sampling as a baseline.
vanilla speculative decoding often suffers from high drafting latency, which can account for over 75\% and 60\% of the total inference time when using Qwen3-1.7B to accelerate Qwen3-14B and Qwen3-32B with draft length 5.
“The draft model with limited capacity struggles to precisely approximate the large-scale target model.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Speculative Sampling. Values are copied from the source paper's tables — verify against the cited paper.
Spiffy beats Speculative Sampling
3.07 vs 1
Speedup · [LLaDA-Base-8B HumanEval]
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative DecodingECHO beats Speculative Sampling
5.25 vs 1.76
Avg. Speedup · [Vicuna-13B]
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency ScenariosDART beats Speculative Sampling
2.77 vs 0.98
Speedup · [Qwen14B Temperature=0]
DART: Diffusion-Inspired Speculative Decoding for Fast LLM InferenceLogitSpec beats Speculative Sampling
3.68 vs 1.87
FLy beats Speculative Sampling
5.34 vs 2.82
Speedup · [L31 405B, Temperature=0]
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchDiffuSpec beats Speculative Sampling
3.08 vs 1.67
Speedup · [training-free block, Qwen2.5-32B target]
DiffuSpec: Unlocking Diffusion Language Models for Speculative DecodingDuoDec beats Speculative Sampling
2.61 vs 1.80
speedup (φ) · [Llama2-7b]
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting+ Reflect Verify beats Speculative Sampling
1.36 vs 1.18
Speed · [Llama3.2-1B-Instruct & Llama3.1-8B-Instruct]
Think Before You Accept: Semantic Reflective Verification for Faster Speculative DecodingCoSpec beats Speculative Sampling
86.97 vs 83.98
Mean · [Qwen3-32B, T=1]
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Collaborative Speculative Decoding (CoSpec)Beyond the Target: From Imitation to Collaboration in Speculative DecodingMay 24, 2026
- ToolSpecToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative DecodingApr 15, 2026
- QuasarQuasar: Quantized Self-Speculative Acceleration for Rapid Inference via Memory-Efficient VerificationMar 2, 2026
- FLy (Training-Free Loosely Speculative Decoding)Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact MatchNov 28, 2025
- Nov 1, 2025
- Oct 30, 2025
- Oct 22, 2025
- Oct 8, 2025
- Group Tree Optimization (GTO)Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative DecodingSep 26, 2025
- Sep 22, 2025