Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
arXiv:2608.08832v1 Announce Type: new Abstract: Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nod