AVQ-Attention: Adaptive Vector-Quantized Attention
arXiv:2607.12789v1 Announce Type: cross Abstract: The O(N^2) complexity of attention over N tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces thi