VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
arXiv:2607.12756v1 Announce Type: new Abstract: Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated