Tutorials
Representation Learning in Diffusion and Flow-based Model: An Application Aspect
arXiv:2608.24068v1 Announce Type: new Abstract: Diffusion models and flow-based models have recently become the dominant paradigms in generative modeling, largely due to their ability to learn rich, m
arXiv:2608.24068v1 Announce Type: new Abstract: Diffusion models and flow-based models have recently become the dominant paradigms in generative modeling, largely due to their ability to learn rich, multi-level visual representations through large-scale training. This creates a bidirectional relationship between generative models and representation learning: improving representation learning enhances generation quality, while the learned representations can be leveraged for broader understanding tasks. This survey systematically explores this interplay with a focus on applications. We propose a three-tier progressive framework that organizes existing works from three perspectives: using representation learning to improve generative capabilities, exploiting generative models to extract representations for perception tasks, and ultimately moving toward general-purpose unified applications. We systematically categorize representative methods across a wide range of downstream tasks, including image classification, dense visual prediction, instance-level perception, and annotation-scarce scenarios. By providing a unified taxonomy and identifying key challenges, this survey aims to clarify the underlying logic of current research and suggest promising directions for future exploration. We hope this work can serve as a valuable reference for researchers interested in harnessing the representation power of generative models for applications beyond generation.
Related
- MeDUET: Disentangled Unified Pretraining for 3D Medical Image Synthesis and Analysis
- Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions
- ARAPDiffusion: ARAP Regularization for Diffusion-Based Deformable Shape Space Learning
- Making Every Step Count: Spatio-Temporal Information Allocation for Imaging Inverse Problems
- Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces
Source: arXiv cs.CV | 2026-08-26