The Key to Going Linear: Analysis-Driven Transformer Linearization
arXiv:2607.07706v1 Announce Type: new Abstract: The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exi