Living systematic review
Speculative decoding
Speeding up autoregressive LLM generation by drafting tokens cheaply and verifying them in parallel.
182 papers333 critique receipts1,849 benchmark resultsupdated Jun 18, 2026
Most-superseded baselines
Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.
- 1EAGLE-3
EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
13 critique · 28 beaten on benchmarks
- 2EAGLE-2
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
10 critique · 21 beaten on benchmarks
- 3EAGLEin EAGLE-2
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
19 critique · 12 beaten on benchmarks
- 4Medusain EAGLE-2
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
19 critique · 9 beaten on benchmarks
- 5Lookahead
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
8 critique · 13 beaten on benchmarks
- 6PLDin Lookahead
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
7 critique · 13 beaten on benchmarks
- 7RESTin Lookahead
REST: Retrieval-Based Speculative Decoding
6 critique · 9 beaten on benchmarks
- 8SpecInfer
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
10 critique · 4 beaten on benchmarks
- 9Speculative Samplingin Lookahead
Speculative Sampling for Parametric Temporal Point Processes
2 critique · 10 beaten on benchmarks
- 10DFlashin EAGLE-3
DFlash: Block Diffusion for Flash Speculative Decoding
4 critique · 6 beaten on benchmarks
- 11LayerSkipin SpecInfer
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
7 critique · 3 beaten on benchmarks
- 12FR-Spec
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
4 critique · 5 beaten on benchmarks
The competition
Methods that fight on the same benchmarks cluster into distinct sub-problems.
Lookahead47 methods
Lookahead · PLD · REST · Speculative Sampling · Token Recycling · PEARL
RSD8 methods
RSD · EARS · EASD · trained evaluation models in speculative decoding · From Tokens to Steps · Entropy-Aware Speculative Decoding (EASD)
SpecReason7 methods
SpecReason · LLM-as-a-judge for sequence-level verification · token-level speculative decoding · SpecThinking · SpecSampling · Lookahead Reasoning
The frontier
Recent methods not yet superseded in the knowledge base.
- Jun 4, 2026
- Jun 3, 2026
- Jun 2, 2026
- May 31, 2026
- May 30, 2026
- May 29, 2026
- May 28, 2026
- May 28, 2026
- May 28, 2026
- May 24, 2026
- May 19, 2026
- May 19, 2026