Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
arXiv:2607.23125v1 Announce Type: new Abstract: Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods