Long-context / context-window extension
DCA
Superseded baseline#16 of 53 most-superseded
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 2 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites DCA as a baseline.
The two training-free length extrapolation baselines, Dual Chunk Attention and Positional Interpolation, shown in Table~tab:qwen_math_extrapolation, achieve accuracies close to zero, demonstrating astonishingly poor performance.
“Fixed group sizes are used for position mapping, regardless of varying input lengths. This lack of adaptability prevents the optimal utilization of well-trained short-range positions.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating DCA. Values are copied from the source paper's tables — verify against the cited paper.
LaMPE beats DCA
59.48 vs 15.96
Avg. · [128K extrapolation]
LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without TrainingDoPE-by-Gaussian beats DCA
0.393 vs 0
Needle Insert (8K) · [8K context Many-shot]
DoPE: Denoising Rotary Position Embedding
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.