Research
JoLT: Joint Latent Trajectories for Context-Guided High-Resolution Tiled Generation
arXiv:2608.15395v1 Announce Type: new Abstract: Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current li
arXiv:2608.15395v1 Announce Type: new Abstract: Although text-to-image generative models produce impressive results, they struggle to generate densely detailed, high-resolution (HR) images. Current literature addresses this issue with a low-to-high-resolution approach. First, a low-resolution (LR) image is generated. Then, an upsampled version is generated using the LR image as an additional cue. In this paper, we present Joint Latent Trajectories (JoLT). To generate an image, JoLT uses two streams that jointly denoise LR and HR latent images at each sampling step. The LR latent controls the overall layout, while the HR latent controls the details. We interconnect both branches to jointly integrate their information. We extensively validate our method, demonstrating its advantages over competing baselines. The resulting images are not only richly detailed but also visually pleasing, opening new avenues for artistic creation.
Source: arXiv cs.CV | 2026-08-18