LGFeb 18

A Residual-Aware Theory of Position Bias in Transformers

Hanna Herasimchyk, Robin Labryga, Tomislav Prusina, Sören Laue

arXiv:2602.16837v14.93 citationsh-index: 2

Originality Highly original

AI Analysis

This provides a foundational architectural explanation for position bias in Transformers, addressing a key issue in NLP and AI model behavior.

The paper tackled the discrepancy between theoretical predictions of attention collapse in Transformers and practical observations, showing that residual connections prevent collapse and cause a U-shaped position bias, explaining the Lost-in-the-Middle phenomenon.

Transformer models systematically favor certain token positions, yet the architectural origins of this position bias remain poorly understood. Under causal masking at infinite depth, prior theoretical analyses of attention rollout predict an inevitable collapse of attention onto the first token. Such collapse, however, does not occur in practice. We resolve this discrepancy with a residual-aware theory of cumulative attention rollout. By incorporating residual connections, we show that this architectural component prevents collapse under realistic conditions. At finite depth, we prove that causal Transformers induce a U-shaped position bias, with attention concentrating on early and late tokens. This result provides a principled architectural explanation for the Lost-in-the-Middle phenomenon.

View on arXiv PDF

Similar