Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
DGX agentarXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin