OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
DGX agentarXiv:2605.29657v1 Announce Type: cross Abstract: Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and