Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
DGX agentarXiv:2602.10743v2 Announce Type: replace Abstract: State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but