Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
arXiv:2511.02043v4 Announce Type: replace Abstract: Attention is a fundamental building block of large language models (LLMs), so there have been many efforts to implement it efficiently. For example,