DiffPrune: differentiable information throttling for token pruning in vision-language models
DGX agentarXiv:2608.01985v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score th