What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
DGX agentarXiv:2608.00013v1 Announce Type: new Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fu