LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
DGX agentarXiv:2604.11789v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks req