Long-context / context-window extension

KERPLE

KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation

Superseded baseline#12 of 53 most-superseded · first seen May 20, 2022

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 3 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites KERPLE as a baseline.

the learned static positional encoding (such as Kerple and FIRE) is an average optimal solution across all training samples. Consequently, while they might be generally effective, they are inherently suboptimal for any specific instance.
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
However, this incorporation of additional trainable parameters results in diminished training velocities.
MEP: Multiple Kernel Learning Enhancing Relative Positional Encoding Length Extrapolation

Beaten on benchmarks

Head-to-head results where a newer method reports beating KERPLE. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.