Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
DGX agentarXiv:2607.20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval rema