Disentanglement-Based Equivariant Learning for Compositional VQA
DGX agentarXiv:2606.02168v1 Announce Type: new Abstract: Compositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of pre