Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
DGX agentarXiv:2607.07953v1 Announce Type: cross Abstract: Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at