Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling
DGX agentarXiv:2606.10435v1 Announce Type: cross Abstract: Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales qua