Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution
arXiv:2607.26596v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified tran