HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference
arXiv:2607.04302v1 Announce Type: cross Abstract: We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on