ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
DGX agentarXiv:2607.02907v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step