TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition
DGX agentarXiv:2606.25478v1 Announce Type: new Abstract: Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generaliza