Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
DGX agentarXiv:2606.31903v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Existing acce