Long-context / context-window extension

NTK-aware

Superseded baseline#6 of 53 most-superseded

Superseded — cited as a baseline and beaten by newer methods

4 papers critique it · 6 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites NTK-aware as a baseline.

rescaling factors derived from previous methods often fall short of achieving the effective target context length.
LongRoPE2: Near-Lossless LLM Context Window Scaling
PE and NTK experience repetition issues, resulting in lower NoRepeat Score.
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
However, these static approaches do not account for the distinctive spectral progression of the diffusion process, where low-frequency structures are generated in the first sampling steps, while high-frequency details are resolved later
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
methods like NTK, Dyn-NTK, and YaRN suffer from attention logit outliers due to their positional embedding interpolations
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)

Beaten on benchmarks

Head-to-head results where a newer method reports beating NTK-aware. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.