Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs
arXiv:2606.08511v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) face a significant inference bottleneck due to the quadratic computational cost of self-attention over long vis