Speculative decoding

EdgeLLM

EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting

Superseded baseline#52 of 151 most-superseded · first seen Jun 22, 2024

Superseded — cited as a baseline and beaten by newer methods

1 papers critique it · 1 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites EdgeLLM as a baseline.

EdgeLLM xu2024edgellm employs local parallel tree generation but imposes heavy memory and computing burdens on limited hardware.
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference

Beaten on benchmarks

Head-to-head results where a newer method reports beating EdgeLLM. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.