Long-context / context-window extension

Activation Beacon

Long Context Compression with Activation Beacon

Superseded baseline#10 of 53 most-superseded · first seen Jan 7, 2024

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 4 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Activation Beacon as a baseline.

However, they need a copy of multi-head attention, which amounts to approximately 2B for 7B models.
FreqKV: Frequency Domain Key-Value Compression for Efficient Context Window Extension
the compressed length still grows linearly with the original context length. This fails to fundamentally alter the asymptotic order of spatiotemporal complexity and can only improve efficiency by reducing constant factors
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling

Beaten on benchmarks

Head-to-head results where a newer method reports beating Activation Beacon. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.