Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
DGX agentarXiv:2607.02593v1 Announce Type: cross Abstract: While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying sole