What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
DGX agentarXiv:2607.23271v1 Announce Type: cross Abstract: Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., 'a dog' vs. 'not a dog') to nearly identical embeddings, rend