Long-context / context-window extension
FIRE
Superseded baseline#24 of 53 most-superseded
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites FIRE as a baseline.
Although FIRE utilizes MLPs to learn positional embeddings, these embeddings remain fixed across different tasks once the training is completed.
“However, as our experiments demonstrate, this behavior was not beneficial in our settings, leading to inferior performance compared to ALiBi and Kerple.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating FIRE. Values are copied from the source paper's tables — verify against the cited paper.
DAPE-Kerple beats FIRE
3.8642 vs 308.6173
perplexity (mean) · [training_length_512_eval_8192]
DAPE: Data-Adaptive Positional Encoding for Length Extrapolation
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Mask Prior Suppression and Monotonic RoPE ScalingMitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language ModelsMay 14, 2026
- Apr 1, 2026
- C^2RoPEC^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models ReasoningFeb 11, 2026
- Imaginary Extension of Rotary Position EmbeddingsBeyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMsDec 8, 2025
- Nov 21, 2025