Text-Conditional JEPA for Learning Semantically Rich Visual Representations
DGX agentarXiv:2605.03245v1 Announce Type: cross Abstract: Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature pre