Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders
DGX agentarXiv:2606.21197v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioni