Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation
DGX agentarXiv:2508.05008v2 Announce Type: replace Abstract: Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their ap