Tutorials
Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces
arXiv:2608.13455v1 Announce Type: new Abstract: Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation r
arXiv:2608.13455v1 Announce Type: new Abstract: Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation models within the representation tokenizer framework and examine whether demographic and clinical information encoded in latent representations from foundation models is preserved during synthetic image generation. We show that generated representations and images faithfully inherit phenotype information when evaluated within their originating foundation models, consistently outperforming conventional latent diffusion on multiple downstream prediction tasks. However, these gains largely disappear when evaluated using classifiers trained on real images, revealing a previously uncharacterised synthetic-to-real representation gap. These findings demonstrate that foundation-model latent spaces provide a powerful substrate for controllable retinal synthesis while highlighting the need to better align synthetic representations with real-image distributions.
Related
- TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model
- Constructing VAE Latent Spaces with Prescribed Topology
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
Source: arXiv cs.CV | 2026-08-14