G^2TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
DGX agentarXiv:2605.12309v1 Announce Type: new Abstract: The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. I