MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
DGX agentarXiv:2605.26004v1 Announce Type: cross Abstract: Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with