Research
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in lever
Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in leveraging high-quality pre-trained visual representations—either by aligning them inside VAEs or directly within the generative model. However, adapting such representations remains challenging due to fundamental mismatches between understanding-oriented features and generation-friendly latent spaces. Representation encoders benefit from high-dimensional latents that capture diverse hypotheses for…
Related
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Diffusing in the Right Space: A Systematic Study of Latent Diffusability
Source: Apple ML Research | 2026-07-15